Wholestory

Last Updated: September 19, 2026

Real-World Adoption: What AI Actually Does

% reporting AI use, by what was measured

Share of businesses
Share of adults / workers
Self-selected / leadership samples
When measured

devices (cumulative)

FDA-authorized AI/ML-enabled medical devices (cumulative)
Year
Measured adoption rates diverge by tens of points depending on who is asked and how; the second view tracks AI systems cleared for clinical use.

Two surveys, three days apart, on the same question. EY asked 202 senior AI executives at US firms above $1 billion in revenue: 91 per cent say their organisation uses agentic AI, 98 per cent have formal AI governance policies. The University of Konstanz asked 1,105 employees: the share using AI at work rose from 35 to 38 per cent in a year, and only 55 per cent of those say their most-used tool was officially introduced by their employer. They are not contradictory — one asks about organisational deployment, the other about personal use. But they are the distance between a boardroom and a desk, and only one of the two sells the governance consulting its findings argue for. Where AI is used, it is used unevenly: 49 per cent of office workers against 25 per cent in manual occupations, 56 per cent of the highly educated against 21 per cent. In small organisations 11 per cent have had training. EY’s own executives report governance they also report not following: 98 per cent have policies, 47 per cent say their organisation has previously not applied them, and 26 per cent cannot detect unauthorised AI agents inside their own walls.

The Whole Story

Ask how much AI is being used and the answer arrives twice, in two different registers. Companies deploying it describe transformation; the measurements that exist describe something narrower, slower and much harder to pin down. The distance between those two accounts is not a matter of opinion — it is a matter of what anyone has actually counted, under what conditions, and with what at stake if the number is wrong. Evidence gathered under liability, measured in controlled studies, or administered by a regulator says one thing. Vendor case studies and surveys of intention say another, and they are the ones most often quoted.

Part of the gap is arithmetic. AI adoption has no single rate, because the rate depends entirely on whom you count. Ask American firms whether they use AI to produce goods and services and roughly one in five say yes; ask American workers whether they have used it on the job and the figure is around 55 percent; ask senior executives or a vendor's own client base and it climbs toward four-fifths. All of these are honest numbers about different populations, and quoting one without its denominator is how a modest technology becomes a revolution in a headline. Adoption also diverges sharply from use: a health system can put an AI scribe in front of 4,000 clinicians in four months and still find it used in only 70 percent of eligible encounters by those who have taken it up. And what predicts adoption turns out not to be technical capability at all — an economic analysis of German workers found that exposure, the share of a job AI could in principle do, explains little, while comparative advantage, AI's output per dollar of cost against a worker's output per dollar of pay, explains most of it. Accountants, heavily exposed, adopt little; primary-school teachers, less exposed, adopt readily.

The other part of the gap is that measured results are genuinely mixed, and have been from the start. IBM's Watson promised to transform cancer care and did not, which set the cautionary precedent. Then AlphaFold solved protein-structure prediction outright, released some 200 million structures, and won a Nobel Prize in Chemistry — proof that the technology can produce results of the first rank in a domain with a clean answer key. Between those poles sit the workplace studies. Controlled trials found support agents 14 percent more productive with the largest gains going to novices, and consultants helped substantially on some tasks and actively harmed on others — the 'jagged frontier'. An MIT survey found 95 percent of enterprise generative-AI pilots showing no measurable return. In medicine the pattern has sharpened rather than resolved: a randomized trial of AI in mammography screening cut interval cancers, while a randomized trial of a language model advising clinicians in Kenyan primary care improved their documentation and diagnoses markedly and left patient outcomes statistically unchanged.

What is unresolved is the step from the task to the total. There is now reasonable evidence that AI makes particular pieces of work faster or better, and very little that those gains aggregate into outcomes for patients, firms or economies — a gap the UN's scientific panel on AI has stated plainly, and one the Kenyan trial illustrates precisely, since its authors calculate that detecting a modest patient-level effect would require more than 100,000 patients. Two structural obstacles keep it open. Almost every large usage dataset belongs to a company selling the product, and national statistics were not built to see this. Meanwhile the checking that would settle it is often not happening at all: when a state auditor examined a 64-institution university system running AI on clinical notes and dropout risk, it found that none of the campuses it sampled had any procedure for testing whether the outputs were accurate. The strongest claim the record currently supports is a modest one — AI is widely used, unevenly, for a narrow band of tasks, with documented wins in a few domains and documented failures in others, and the question of what it adds up to remains open because measuring it is expensive and few of the people deploying it are trying.

Continue Reading →

Observation

91 Per Cent, or 38 Per Cent? It Depends Who You Ask, and Who Is Asking.

Two surveys three days apart. EY asked 202 senior AI executives at US firms above $1 billion in revenue: 91 per cent say their organisation uses agentic AI, 98 per cent have formal AI governance policies. The University of Konstanz asked 1,105 employees: the share using AI at work rose from 35 to 38 per cent in a year, and only 55 per cent of those say their most-used tool was officially introduced by their employer. The two are not contradictory — one asks executives about organisational deployment, the other asks workers about their own hands. But they are the distance between a boardroom and a desk. And the direction of interest runs one way: EY sells the governance assurance its findings argue for. The university does not.

Executives Report Governance They Also Report Not Following

In EY’s survey of 202 senior AI executives, 98 per cent say their organisation has formal AI governance policies — and 47 per cent say it has previously not applied them. Among those using agentic AI, 49 per cent say their governance framework has not been updated for it, and 26 per cent say they cannot detect unauthorised AI agents operating inside their own organisation. Thirty-six per cent report an AI incident or failure causing materially negative impact, including data loss and operational disruption. All of it is what executives say, to a firm that sells the assurance services the findings argue they need.

Half the Office Uses It. A Quarter of the Factory Does.

The Konstanz study finds workplace AI use splitting along familiar lines: 49 per cent among office and knowledge workers against 25 per cent in production and manual occupations; 56 per cent among employees with high educational attainment against 21 per cent with lower. On perception, 40 per cent expect AI and automation to harm the labour market over the next decade — but only 17 per cent fear for their own job, and 12 per cent think AI has already cut jobs at their own employer. A survey of 1,105 employees, self-reported throughout.

Counting AI users by what they pay for: about a quarter of firms, and a 42-point sector gap

Instead of asking firms whether they use AI, Ramp and Revelio Labs watched them pay for it, linking corporate-card spending to workforce records for 21,559 US companies. An adopter spends at least $100 a month with AI vendors for three straight months: 5,633 firms, about a quarter, by December 2025. The spread is the finding — 53.7 percent in information against 11.3 in arts and entertainment, 41.5 percent at 250-999 workers against 14.3 below ten. Adopters were already faster-growing, more engineering-heavy and three times as likely to be venture-backed. Not a national estimate, but a rate nobody self-reported.

Two Federal Reserve studies: AI use is broad but shallow, and firms retrain rather than fire

On the same day, two Federal Reserve studies converged on a picture of AI use that is broad but shallow. The St. Louis Fed's Real-Time Population Survey produced the first nationally representative task-level adoption measures, from about 14,000 workers: at least one in five workers uses AI in more than 80 percent of occupations, but fewer than 3 percent of individual job tasks reach 50 percent adoption, and none reach 70. Its sharper point is that exposure scores predict which occupations adopt, not which individuals do — learning from experience, not demographics, decides that. The New York Fed's regional business surveys found 61 percent of service firms and 51 percent of manufacturers now using AI, up from 40 and 26 percent a year earlier; yet three-quarters call the investment minimal-to-modest, the median adopting firm has only 17 percent of workers using it, and just 4 percent of service firms cut staff because of AI while 13 percent hired more people to run it. Both are Fed research rather than peer-reviewed, but independent of any vendor.

Enterprise AI spending starts to drift toward open-source and Chinese models

Corporate-spend data caught the first measurable drift in US enterprise buying beyond American proprietary labs. Ramp's AI Index, which tracks token and subscription spend across its business-card customers, found the share of AI-spending firms paying for model-serving platforms — the route to open-source and Chinese-developed models — rose to 6.1 percent in July from 4.5 percent in January. American labs still dominate: Anthropic led at 43.5 percent of businesses, OpenAI 39.7, xAI 4, and every first-time buyer still starts with a US lab. The drift is confined to existing advanced spenders fine-tuning cheaper open models on their own data. Ramp measures customers' actual spend, not intentions, though its sample skews slightly tech-heavy; Fortune reported the shift from the same data.

The Federal Statistical System Starts Using Anthropic's and Microsoft's Own Telemetry

BLS has classified every occupation it projects — 831 of them — into four categories of relative AI exposure, to sit alongside its 2025–35 projections. The method is the news. Three of the five inputs are academic exposure scores; the other two are vendors' own product data: Anthropic's, mapping Claude.ai conversations and API traffic onto O*NET tasks, and Microsoft's, mapping observed Copilot use onto work activities. Two companies' telemetry is now an input to an official US occupational classification. BLS is unusually careful about what it has not done: the categories are relative, not absolute; they do not separate automation from augmentation; and exposure implies nothing about job loss, wages, or productivity. Of 4,155 occupation-source pairs, 211 were imputed, affecting 75 occupations. The published table carries employment, wage, education and openings figures for each occupation alongside its category.

22.4 Percent of American Businesses, and Almost None of the Ones That Dig Things Up

The Census Bureau's Business Trends and Outlook Survey, collected 10–23 August, puts AI use at 22.4 percent of US employer businesses, with 25.9 percent expecting to use it within six months. The interesting numbers are the spread. By industry it runs from 46 percent in information and 41.9 in professional and technical services down to 10.2 percent in transportation and warehousing, 8.6 in accommodation and food service and 8.2 in mining. By state, Washington DC leads at 33.7 percent and Utah at 30.3, while California — the industry's home — manages 23.6, barely above the national figure, and New York sits near the bottom at 16.7. Census publishes these as interactive charts rather than text, so the figures here are read from a secondary transcription of its dashboard.

The Government Cannot Yet Tell Whether AI Is Changing the Labour Market

A Washington Center for Equitable Growth brief — the first of three mapping the federal statistical system — concludes that existing federal data cannot show conclusively whether or how AI is reshaping American work. The reasons are specific. Data on unemployment, wages and job availability exist but are fragmented, often untimely, and hard to link across sources. Firm-level adoption is asked about in binary yes-or-no questions, inconsistently between surveys, and never connected to what happens to those firms' workers. The unemployment rate itself comes from a household survey that excludes people who stopped looking and the underemployed, cannot easily be narrowed to an occupation, and has seen its response rate fall for a decade. Weekly jobless claims are timely but cover only those who file — not independent contractors, not the long-unemployed, not those whose benefits ran out — and their demographic detail varies by state.

A Review of a Hundred Studies Finds the Answer Depends on the Question

A working paper surveying more than 100 studies of firm AI use, summarised by WorkRise, finds that results turn on how researchers measure AI at all: by AI-related hiring, by investment, or by firms' reported use. What holds across the literature is that firms investing more in AI are larger and more skill-intensive, and their investment is associated with higher sales growth and market value, typically two to three years afterwards. What does not hold is any consistent read on jobs or productivity; effects vary by use and by measure. Recorded from a summary that names no authors.

Census counts the worker side: 55 percent say they have used AI on the job

The Census Bureau published a federal measure of AI use counted at the level of the worker rather than the firm, from the March 2026 wave of its Household Trends and Outlook Pulse Survey. About 55 percent of U.S. workers said they had used AI at work for at least one of eleven listed tasks — searching for information (37 percent), writing communications and documentation (32 percent), generating ideas (32 percent), summarizing or translating (31 percent) and administrative work (27 percent) leading the list, with coding, customer support and logistics well behind. The figure sits far above the roughly one in five American firms the Bureau's business survey counts as AI users, and the gap is not a contradiction: one question counts companies, the other counts people, and the distance between the two answers is most of what the phrase 'AI adoption' conceals. Frequency thins the number again. Among those who had ever used AI at work, 24 percent had used it every day in the past week, 46 percent on some days, and 30 percent not at all. Use divides sharply by education — 75 percent of AI-using workers with a bachelor's degree or higher had used it in the past week against 59 percent of those with a high-school diploma or less — while the sex gap is one of intensity rather than participation: 30 percent of male AI users used it daily against 17 percent of female users, though roughly equal shares of each had not used it at all that week. On what it delivered, the survey asked workers to estimate how many additional hours they would have needed without it. Thirty-one percent said AI saved them one to two hours and 25 percent less than an hour; 15 percent claimed three to four hours and 15 percent more than four. Thirteen percent said it saved them nothing or cost them time. Those are self-assessed counterfactuals rather than measured savings — no baseline, no control — and belong in a different evidence class from a controlled trial. The underlying measure is a probability household sample with comparisons tested at the 90 percent confidence level, but the Bureau's account gives no fieldwork dates beyond the month, no sample size and no margins of error.

An auditor asked a university system whether it checks its AI. None of the four campuses did.

The New York State Comptroller published a performance audit of the State University of New York — 64 institutions, the largest public higher-education system in the country — examining how it governs the AI it has already deployed. Its sharpest finding is a single sentence: at the four campuses sampled, University at Albany, Stony Brook, Upstate Medical and Onondaga Community College, 'none of the campuses required or implemented specific procedures to test their AI systems to evaluate whether outputs were accurate or biased.' The audit also puts on the record what is actually running. SUNY campuses use AI to transcribe clinical notes during patient visits, to read licence plates and identify parking violations, to monitor campus data for students at risk of dropping out, and to mask personally identifiable information in video evidence. Auditors found the system's central administration had no effective AI governance framework, no standard definition of AI, and no documented policies for developing or using these systems, with practice varying widely between campuses. It is a narrow document — one state, one system, covering January 2019 to October 2025 — and its severity framing is an elected auditor's. But it is one of the few accounts of AI deployment written by someone with the authority to go and check, rather than by the organisation doing the deploying. Its answer is not that the technology failed. It is that at an institution already using AI on patient notes and dropout risk, nobody had been measuring whether its outputs were right.

Why workers actually adopt AI: comparative advantage, not raw capability

A study by Ilse Lindenlaub, Ryungha Oh, María Alejandra Rodríguez and Laura Veldkamp set out to explain the gap between what AI can do and whether workers use it, combining a nationally representative survey of actual AI use by German workers — linked to official worker and establishment records including wages — with an economic model of adoption. Its central finding is that technical 'exposure', how much of a job AI could in principle do, is a poor predictor of whether an individual worker adopts it. What decides adoption is comparative advantage: AI's output per dollar of user cost against a worker's output per dollar of pay. Where user costs are high — verification of the AI's work, workflow integration, regulatory compliance, data privacy — adoption stalls even at high exposure, which is why accountants, heavily exposed on paper, adopt little, while primary-school teachers, less exposed but facing low user costs, adopt robustly. Looking ahead, the authors project AI adoption nearly doubling to about 80 percent over three years — meaning most workers using AI for at least one task — driven mainly by falling user costs rather than rising capability, and spreading to more workers without much widening the number of tasks each uses it for. The working paper puts numbers to how much better that explanation performs: its occupation-level comparative-advantage index accounts for almost 60 percent of the variation in observed adoption between occupations, against 14 percent for a model built on exposure alone, and the two approaches disagree substantially about roughly 30 percent of workers. The evidence stands out on this page for being independent of any vendor and grounded in administrative data. The three-year projection is the authors' as relayed in Yale Economics' summary of the paper; the working paper's public abstract states the explanatory figures above but not the forecast, and its full text is not openly available.

Science cannot yet say with confidence whether the task-level of AI productivity gains will aggregate to economy-wide gains.?

The UN's Independent International Scientific Panel on AI — 40 experts appointed by the General Assembly, co-chaired by Yoshua Bengio and Maria Ressa — set out the evidence gap in its first preliminary report. The available research measures cost reductions in tasks that already exist far better than it measures AI's contribution to new goods, services and markets, so the step from a saved hour to a larger economy remains unevidenced. The panel also identified why the gap persists: there is no common standard for privacy-preserving analysis of AI usage by the companies that produce it, and the economic effects will stay hard to measure unless national statistical systems are updated and given measurement access. That is the reason nearly every adoption figure in circulation is supplied by a vendor. The report notes it does not represent the views of the United Nations.

Context: Still unverified, and now sharply illustrated: a registered randomized trial of AI decision support in Kenyan primary care found clinicians' diagnoses and notes significantly improved while patient outcomes did not move (odds ratio 0.77, P = 0.13). Its authors estimate that detecting even a modest patient-level effect would require more than 100,000 patients.

The hardest test yet of AI in the clinic: better notes, no better outcomes

A cluster-randomized controlled trial of a large-language-model assistant in Kenyan primary care, published in Nature Medicine, is the most demanding evidence yet gathered on whether AI help for clinicians reaches patients — and it found that it did not. Across 16 clinics run by the operator Penda Health in Nairobi and Kiambu counties, 103 clinical officers were randomly assigned to an electronic medical record with or without 'AI Consult', a tool that read what they typed and flagged concerns green, yellow or red. Some 9,691 patients were enrolled between April and July 2025. The trial's primary measure was treatment failure within 14 days — a patient returning with unresolved symptoms, needing escalated care, or suffering a related adverse event. It occurred in 102 of 4,693 patients (2.2 percent) where the AI was switched on and 94 of 4,654 (2.0 percent) where it was off: an adjusted odds ratio of 0.77, with a confidence interval from 0.55 to 1.08 that comfortably spans no effect. What did move was the quality of the clinician's work. A blinded panel of six Kenyan family physicians reviewing 2,000 encounters found the assisted clinicians markedly more likely to record an appropriate diagnosis, a comprehensive note and an appropriate treatment plan — odds ratios near 1.7 on all three, each significant beyond the 0.001 level. Yet none of the trial's specific clinical targets shifted: antibiotic and antimalarial prescribing, hypertension diagnosis and treatment, child malnutrition referral all showed no difference, and the one statistically significant result ran the other way, with patients in the AI arm slightly less likely to be flagged as at risk of type 2 diabetes. The result cannot rule out a small benefit — the trial was built to detect a halving of treatment failure and saw nothing of that size — but it rules out a large one, and the authors put a price on settling the question: detecting a modest difference in outcomes this rare would take more than 100,000 patients, and they chose not to continue. Running the model cost about four U.S. cents per consultation. Its evidence level is as high as deployment evidence gets: registered with the Pan-African Clinical Trials Registry before it began, with its protocol and analysis plan deposited publicly a year before the results, reported against the CONSORT-AI standard, and overseen by an independent safety board. The qualification is that the tool was built by the same company whose clinics hosted the trial, and the trial's own paperwork discloses what a summary would not — 921 protocol deviations, among them a temporary configuration error that briefly exposed some control-arm clinicians to the AI. The University of Birmingham, whose researchers led the work, announced it under the headline 'AI clinical support tool improved clinician decisions in real-world primary care trial', disclosing the null patient result in the subheading beneath; the specific prescribing and diagnosis findings, including the diabetes result favouring the control group, appear only in the paper.

Cleveland Clinic scales an AI scribe to 4,000 clinicians in four months — and shows what 'use' really means

A deployment report in npj Health Systems documented how the Cleveland Clinic, partnering with the vendor Ambience Healthcare, onboarded more than 4,000 ambulatory-care clinicians — about 80 percent of its ambulatory workforce — to an ambient AI scribe in under four months. Twelve months on, more than 4,800 clinicians had used the tool across over 3.5 million encounters, at 70 percent encounter-level utilization among established users, with a Net Promoter Score of 60, a CSAT score of 96.6 percent, and 60 percent of users agreeing it had increased their likelihood of remaining in practice. The report's most useful contribution is conceptual: it insists on the difference between adoption — whether a clinician has ever used the tool, reported at 20 to 42 percent across prior literature — and utilization, the more demanding share of eligible encounters in which it is actually used. It reads as a deployment playbook (governance, phased training, staffed 24/7 support, real-time feedback), and its provenance sets its evidence level: a single health system's experience, co-authored with the vendor whose product it deploys, with no control arm and self-reported satisfaction measures. It is recorded here as measured enterprise-scale usage in a clinical setting, with that provenance labeled.

An independent RCT: AI image-triage speeds Brazil's textbook review without losing quality

A randomized controlled study in Scientific Reports tested an AI component inside Brazil's national textbook program (PNLD), a government process that can take at least two years and involves hundreds of analysts. In a parallel two-arm design with 1:1 randomization across 76 textbook analysts, the AI — a convolutional neural network classifying textbook images as sharp, defocused-blurred or motion-blurred — significantly increased assessment productivity while preserving quality: the experimental group assessed substantially more images than the control, with a modest improvement in analysts' ability to distinguish defect categories. What makes it notable against the rest of this record is its provenance: the authors declare no competing interests, making it an independent, non-vendor randomized evaluation of an AI decision-support tool in a real public process — rarer here than the maker-authored medical trials, which enter with their conflicts labeled. Its limits are scope and endpoint: the task is a narrow image-quality triage, and the study was not registered in a clinical-trial registry because it is an educational-policy rather than a health intervention. It enters at RCT evidence level.

OpenAI's enterprise data: work turns 'agentic', and the gap between lead firms and the rest widens

Alongside its consumer country data, OpenAI published two enterprise-usage reports — Enterprise Signals and a working paper, 'How Organizations Use AI: Evidence from ChatGPT' — drawn from administrative usage across its enterprise customer base rather than from a survey. The picture is of use deepening unevenly. Work is becoming more 'agentic': by June, Codex generated 64 percent of the combined Codex-and-ChatGPT output tokens among enterprise customers, a shift from answering questions toward carrying out multi-step tasks. And the spread between firms is widening — 'frontier' firms, the top tenth by usage each month, now generate 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January, a gap that holds across industries and company sizes. Agentic use is spreading beyond engineering: weekly-active enterprise Codex users grew since February by 108 times in legal, 41 times in sales and recruiting and 26 times in marketing, against 5 times in engineering. Contrary to surveys that find leaders using AI most, the administrative data show early-career employees sending 13 more messages a week than executives six months after adoption. The companion paper reports that enterprise adopters among U.S. public companies had stronger financials than non-adopters — more assets, employees and R&D — but that is association, not a liability-disclosed outcome, and OpenAI cautions that 'access alone may not be enough.' Read it with the same provenance as the consumer release: one vendor counting its own product's traffic, enterprise accounts only, output-token and message counts as proxies for 'depth', and every use-case label assigned by a model rather than a human. Its value is that it fills the enterprise blind spot that the Google ATLAS, IMF and OpenAI consumer analyses all excluded by construction.

OpenAI's first country-by-country usage data: more "doing" at work, multimedia rising, the global gap narrowing

OpenAI published its first country-level dataset of measured ChatGPT usage, drawn from classified individual-account messages (Free, Go, Plus and Pro) across 144 countries and released through its OpenAI Signals hub. Because it watches what people did rather than asking what they think they do, it belongs with the observational usage counts rather than the intention surveys. The findings: at work people are more than twice as likely to use ChatGPT to complete or produce something — editing, coding, analysis — than outside it, where information-seeking still dominates; multimedia has climbed to 7.8 percent of messages worldwide since the April image tools shipped, and above one in ten in Brazil and Colombia; the share of messages from users 35 and older rose year-over-year in almost every country, by more than ten percentage points in France and Czechia; and per-capita usage in parts of Latin America, Africa and Oceania (Peru, Uruguay and Costa Rica rising fastest) grew quicker than in the established early-adopter markets, narrowing the global gap. Read it with its provenance in view. It is one vendor's count of its own product's traffic, so a person or country using a rival tool registers as a non-adopter, and the dataset deliberately excludes enterprise and organization accounts, the same blind spot the Google ATLAS and IMF analyses flagged. Every use-case and occupation label is assigned by a language model, not a human coder. OpenAI states that "more than 1 billion people" now use ChatGPT — its own figure, not an independently measured one, the same status as the 800-million-weekly count it gave in late 2025.

Vendor-authored trial: AI "continuous care" nudges psychotherapy engagement and symptoms — modestly

A preregistered study in npj Digital Medicine tested whether AI-enabled "continuous care" features embedded in employer-sponsored psychotherapy improved on therapy alone. In a cluster-matched, quasi-experimental design, adults beginning therapy at 25 employers whose plans included the features were compared with participants at 75 matched employers without them. The effects were real but small: access was associated with 5 percent greater session attendance (rate ratio 1.05, 95% CI 1.01–1.10; n=26,208 tracked over seven weeks) and a faster time to a second session, and — on top of the substantial improvement patients got from psychotherapy alone — modest additional reductions in depression and anxiety symptoms (Cohen's d 0.15–0.16; number-needed-to-treat 25; symptom cohort n=5,518 followed up to six months). The conflict of interest is on the face of the paper and shapes how the result is read: all eight authors are employed by and hold equity in Spring Health, the company whose product is under study, the work was funded by Spring Health's own internal resources with no external funding, and the corresponding author additionally reports patents and equity across several health-technology firms. It is a large, real-world, peer-reviewed deployment evaluation — stronger than a press-released claim — but a maker measuring its own feature, at a modest effect size honestly reported, and it enters at that level with its provenance labeled.

Vendor-funded trial: AI-assisted blood-smear reading beats manual microscopy across 15 cell categories

A multicenter, randomized, paired clinical-validation study in npj Digital Medicine — 1,570 peripheral blood smears read across three Chinese centers — found AI-assisted digital morphology outperformed manual light microscopy across all 15 nucleated-cell categories. It improved screening for acute promyelocytic leukemia, a hematologic emergency (abnormal-promyelocyte detection F1 0.79 [95% CI 0.67-0.89] vs 0.66 [0.51-0.77]), improved most erythrocyte categories including teardrop cells (F1 0.31 vs 0.26), gave more accurate platelet estimation (absolute mean deviation from flow cytometry 2.05 vs 14.08 x10^9/L), and raised technician efficiency by roughly 60 percent. The conflict of interest is on the face of the paper and matters to how the result is read: the study was funded by Shenzhen Mindray Bio-Medical Electronics as the regulatory clinical performance evaluation of its own MC-100i digital-morphology analyzer, submitted toward China NMPA approval, and two Mindray employees are among the authors. The disclosure states the sponsor had no role in the study's design, conduct, analysis, interpretation, writing, or publication decision, and that the academic authors had full data access and final responsibility. The evidence is stronger than a press-released vendor claim — peer-reviewed, prospective, randomized-paired and multicenter — but it is a diagnostic-accuracy and efficiency validation of a maker's own device, not an independent patient-outcome trial; it enters at that level, with its provenance labeled.

The ECB counts 38 percent of euro-area firms as advanced AI adopters

An ECB Economic Bulletin box put 38 percent of euro-area firms at an advanced stage of AI adoption — significant or moderate use, by the survey's own scale — the central bank's firmest adoption figure to date for the bloc. The number lands mid-way between the boosterish vendor surveys and the sparse official statistics this record usually has to choose between, and because the ECB ties it to investment expectations, it will feed the bank's own capital-formation forecasts: adoption is now an input to European monetary policy analysis, not just a technology story.

Half of AI unicorns have never published a qualifying scientific paper

A bioRxiv analysis of 317 AI unicorn startups found 52.4 percent produced no qualifying scientific output between 1998 and 2025, only 24 firms — 7.6 percent — produced any highly cited paper, and three companies account for 92 of the 134 such papers among them; unicorn participation in formal scientific literature ran to 0.1 percent of AI publishing in 2025. A preprint, unrefereed, and 'qualifying output' is a definition with edges — but the concentration it measures is stark: the billion-dollar tier of the AI industry overwhelmingly applies science produced elsewhere, with research influence pooled in a handful of firms.

Hospital AI adoption tracks telehealth — and the hospitals not reporting are the ones without it

A national study of 6,173 U.S. acute-care hospitals found telehealth volume the strongest predictor of AI adoption in both clinical and operational domains — digital infrastructure begets digital infrastructure. The measurement gap is the sharper finding: 57 percent of hospitals did not report telehealth volume at all, and those non-reporters account for 91.4 percent of hospitals with no observable AI adoption. Whatever the true rural-urban adoption divide is, the national statistics are computed mostly from the hospitals already digitized enough to be counted.

Compliance is the caboose: firms adopt AI three times faster than their watchdogs

Ethisphere's AI and data report put broad or advanced AI adoption at 67 percent of surveyed organizations — while only 22 percent of the same organizations' ethics and compliance functions had reached equivalent adoption, a 45-point governance lag inside the enterprise. The compliance teams' own stated barriers are the technology's familiar failure modes: 53 percent cite accuracy and hallucination risk, 48 percent confidentiality and data exposure. A certification vendor's survey of its own market, with that caveat — but the shape it describes, deployment running years ahead of internal oversight, matches what this record shows at the regulatory scale.

The small-business AI gender gap is widening, and youngest firms show it most

A second JPMorganChase Institute report from the same transaction series found male-owned small businesses adopting AI at consistently higher rates than female-owned ones, with the gap widening since 2023 rather than closing as costs fell. Among Generation Z founders — the cohort with the least pre-AI habit to unlearn — 20 percent of male-owned firms had adopted by 2025 against 13.9 percent of female-owned firms. Falling prices were supposed to be the equalizer; the series suggests the divide is not primarily a price effect, which matters for every program predicated on cheap tools closing adoption gaps on their own.

Six years of adoption in six months: JPMorgan charts small business AI's collapse in entry costs

JPMorganChase Institute analysis of its own small-business transaction data found the 2025 cohort of small businesses reached 10 percent AI adoption within six months of founding — a threshold the 2019 cohort took over six years to cross. The mechanism is price: firms adopting in 2024 started at a median $20 a month, down 60 percent from the $50 the 2019 cohort paid, and single-service reliance fell from 89 percent of paying firms in 2019 to 72 percent by 2025 as businesses stacked multiple tools. Transaction data measures paying customers, not usage or benefit — but as a cost-of-entry series it is the cleanest evidence yet that AI adoption at the small end is now a default rather than an investment decision.

Google counts 14.7 million of its own AI conversations: wide reach, shallow use, and it skews rich

Google published ATLAS v1.0, an observational count of 14,653,926 de-identified interactions with the Gemini app, Google's AI Mode in Search and the Gemini API over two weeks in April 2026, mapped onto 800-plus occupations and some 4,000 government-catalogued job tasks. Unlike a survey, it watches what people did rather than asking what they think they do — the strongest thing about it. The picture is breadth without depth: some Gemini use appears in 68% of detailed occupations, which the paper says account for 88.4% of employed American civilians, but the typical occupation shows use on only 21% of its tasks, 29% show none at all, and just 3% show it across three-quarters or more. Non-routine analytical work absorbs the traffic — 35% of catalogued tasks, 65% of work interactions — while attempts to hand a task over end to end run below 10%. More than 86% of all interactions happened outside work altogether, spanning activities that fill 98% of Americans' waking hours, with government and civic business appearing almost twenty times more often than the time people actually spend on it and nearly half of medical, legal, financial and government queries arriving outside office hours. The distributional finding matches what the IMF found independently in a rival vendor's data two weeks earlier: a 1% higher occupational median wage goes with more than 2.5% more usage, the conversation-weighted median salary is about $83,000 against a national median some $20,000 lower, and across countries a 1% higher GDP per head goes with 0.9% more use, leaving the least-adopting fifth of countries with 2% of conversations. Read it with three things in view. Every occupational label is assigned by a language model rather than a human coder. It counts Google's products only, so a country or a worker using a rival tool registers as a non-adopter, and it excludes the paid enterprise interfaces entirely — the same blind spot the IMF paper quantifies, for one vendor's business interface alone, at a further $1.14 trillion a year. And it is not peer-reviewed: the two outside academics Google credits with reviewing it, David Autor and Diane Coyle, are both inside the Google-funded programme that produced it. Google says it finds "no evidence... to support the claims that AI is about to cause massive automation and displacement of white-collar work" — but the study also states that it "measures behavioral interactions, not definitive productivity outcomes," and it watches Gemini's users rather than the labour market, so it is not built to see displacement either way.

The IMF puts a price on AI's time savings — $2.7 trillion, 96% of it in rich countries

An IMF working paper by Rachel Yuting Fan and Ha Minh Nguyen valued the time AI currently saves at roughly $2.7 trillion a year, about 3.4% of the combined GDP of its 86 sample countries. The number is an extrapolation and the paper says so: it ranges from $1.6 trillion to $6.1 trillion depending on assumptions about Claude's conversation volume and market share, it values time freed rather than output produced, and the hours-saved figures underneath every dollar are — in the authors' own words — "estimated by Claude, not measured." What is measured, and needs no such scaling, is the distribution. Across five waves of the Anthropic Economic Index (about a million Claude conversations sampled per wave, January 2025 to February 2026, geolocated to more than a hundred countries), the paper's concentration index is positive almost everywhere: 0.49 in the United States, 0.78 in Kenya, 0.84 in India, 0.98 in Tanzania, meaning the savings flow overwhelmingly to the better-paid. In the United States, professionals are 23.2% of employment and take 83.6% of the measured savings, while the bottom 65% of workers by wage have accumulated 10.7%. High-income countries capture 96% of the global total, worth 4.2% of their GDP against roughly 0.6% in middle-income and 0.1% in low-income economies. Two counter-currents run against that: the average AI conversation is getting cheaper work done — the usage-weighted wage fell 5.5% over 13 months as computing and mathematical occupations lost 4.9 points of conversation share to education, sales and office work — and the tilt is easing, with the concentration index falling in 48% of countries in the most recent quarter against 30% in the one before. The strongest predictor of how much value a country captures is not its income but its regulatory readiness; the income effect disappears once readiness is controlled for. The paper's own summary of all this is that "AI is trickling down, but it has not yet reached the bottom." It is staff research, not peer-reviewed; it covers only consumer Claude conversations, excluding enterprise interfaces and every rival platform; and it counts only benefits, netting out neither displacement nor the infrastructure, energy and subscription costs that would be needed for a full welfare account.

A randomized trial finds an AI chatbot no better than a CDC webpage — and shorter-lived

JAMA Network Open published a randomized controlled trial of nearly 1,300 parents in the United States, United Kingdom and Canada, assigned to one of three arms: a conversation with a large-language-model chatbot, government-issued HPV vaccination materials, or no message at all, with each exposure lasting at least three minutes. Intent to vaccinate was measured immediately and again at 15 and 45 days. The chatbot did raise stated intent against no message — but it did not beat the public health materials, and by 45 days its effect had faded while the materials' effect held. Senior author Sharath Chandra Guntuku of the University of Pennsylvania put the result plainly: "AI chatbots are promising, but they should not be assumed to outperform existing tools simply because they are newer or more interactive. A short read of a CDC webpage held up at least as well as a chatbot conversation, and the effect actually lasted longer." Two limits on the evidence: the endpoint is self-reported intention, not observed vaccination, and the trial's effect sizes, confidence intervals and funding are not stated in the summary available here — the journal article itself could not be retrieved.

Five million papers later: AI plus supercomputing tracks with more novel, more-cited science

Researchers at Strasbourg, Turin and Sorbonne Paris Nord published a study in Scientific Reports analysing metadata from more than five million scientific publications across 27 fields from 2000 to 2024, and found that research combining AI with high-performance computing is more likely to introduce novel ideas and to reach top-cited status than conventional research — or than research using either AI or supercomputing alone. It is one of the first attempts to measure whether AI adoption in science is associated with better science, rather than to catalogue individual landmark results. The evidence level matters: this is observational bibliometrics on publication metadata, so it establishes correlation and not causation, and its outcomes are proxies — novelty of ideas and citation rank — rather than validated or replicated discoveries. The paper is open access and peer-reviewed (accepted 20 July 2026), but at the time of writing the publisher was serving only the abstract and front matter, so the effect sizes, methods and the authors' own stated limitations could not be read; the figures here come from the abstract. The same study reports that access to supercomputing and AI expertise is growing more concentrated, in a small number of regions dominated by the United States and China.

Gallup's adoption rate jumps to 47% — because it counts workers, not firms

Gallup reported that 47% of employed U.S. adults say their organization has begun integrating AI, up six points from 41% in the previous quarter, with 52% saying they personally use AI in their role (30% a few times a week or more, 15% daily). The wave surveyed 22,573 employed adults on Gallup's probability-based panel between 6 and 20 May 2026, with a margin of error of ±0.9 points. The number is not comparable to the Census Bureau's roughly 20%: Gallup asks workers whether their employer has adopted AI, while Census asks firms whether they use AI to produce goods or services — the same unit-of-measurement spread the Federal Reserve documented in April across estimates running from 18% to 78%, and Gallup's 47% falls inside that range exactly where the Fed's framework puts a worker-reported figure. What workers use it for is mundane: writing and editing (51% of users), search and research (49%) and general problem-solving (39%), with coding assistance at just 16% — close to the inverse of the specialist mix found among field epidemiologists. Gallup's own two reports, published a day apart from the same panel waves, also show how much the productivity headline depends on where the bar is set: the adoption report says 65% to 77% of users rate AI's effect on their productivity 'extremely or somewhat positive,' while the engagement report says only 17% of frequent users give AI the highest possible rating on the same question. Both figures are self-reported perception, not measured output.

Two-thirds of trainee disease detectives use AI; one in five has been trained on it

A peer-reviewed survey in Eurosurveillance, the ECDC's journal, asked every fellow in the Canadian, U.S. and European Field Epidemiology Training Programs — the pipelines that produce national outbreak investigators — whether they use AI in their work. Of 105 respondents (48% of the 220 invited), 66% said yes, but only 20% had received any AI training and just 47% said their institution had AI policies or guidelines: use has run ahead of both instruction and governance. Among users, ChatGPT dominated at 87% and the overwhelming application was analytical coding assistance at 91%, with 72% using AI at least weekly. Uptake varied sharply by programme — 89% in Europe (32 of 36) against 54% in both the U.S. (30 of 56) and Canada (7 of 13, a cell too small to lean on). Evidence level: self-reported use, not instrumented measurement, and the fieldwork ran 25 November to 5 December 2024 — so these figures describe late 2024, roughly 19 months before publication, and predate most of the model releases since. Reported concerns were accuracy and reproducibility, bias, permissible use and data privacy; the study measured no outcome, error rate or time saved.

More than 800 million people use ChatGPT every week.?

OpenAI CEO Sam Altman said ChatGPT had passed 800 million weekly active users, up from about 500 million in March and 700 million in August 2025, alongside 4 million developers. The figure is self-reported: OpenAI does not publish audited usage metrics and no independent count corroborates the exact number, though outside observers broadly agree ChatGPT is the most-used consumer AI product.

Context: Still unverified: self-reported by OpenAI's CEO on 2025-10-06, and no independent measure of any consumer AI product's user base exists to test it against. The trajectory (500M in March → 700M → 800M by October 2025, 'more than 1 billion' by 2026) has continued on the company's own telling, but the largest independent usage studies (Google's ATLAS, the IMF/Anthropic paper) exclude ChatGPT by construction, and a UN scientific panel notes no provider publishes a comparable cross-platform count.

Anthropic's usage data: AI touches a third of technical tasks, with no unemployment spike yet

Anthropic's Economic Index and labor-market analyses — measured from actual Claude usage and matched to official labor data, though drawn from one lab's self-selected users — found Claude conversations cover about 33% of the tasks in computer and mathematical occupations (against a 94% theoretical ceiling), topping out at 75% for computer programmers. Anthropic reported no systematic rise in unemployment for the most AI-exposed workers through the period studied, though workers aged 22-25 in exposed jobs saw roughly a 14% lower job-finding rate. About 93% of Claude conversations produced a concrete work artifact, and personal (non-work) use rose from about a third of conversations on weekdays toward half on weekends. Real usage measurement, with its motivated-source provenance labeled.

FDA's authorized AI medical devices reach 1,524 — three-quarters in radiology

The FDA's AI-Enabled Medical Device List reached 1,524 authorized AI/ML-enabled devices, of which 1,164 (about 76%) are in radiology. The count has climbed from 64 in 2020 and 950 as of August 2024 — the clearest regulator-administered measure of AI moving into frontline clinical deployment, and the deployment curve this page's instrument tracks. Growth is real but heavily concentrated in imaging.

Census: about one in five U.S. firms now uses AI — but the divide is by size

The Census Bureau's Business Trends and Outlook Survey found overall U.S. business AI use hovering between 17% and 20% (a national rate near 19.8% as of the collection period ending May 3, 2026), the clearest population-scale measurement of enterprise adoption. The gap is by firm size: 37% of firms with at least 250 employees reported using AI versus under 20% of the smallest firms, and use ran highest in the Information (39.7%) and Finance and Insurance (33.9%) sectors. A second AI supplement launched in November 2025 broadened the survey to 15 business functions. Measured adoption is broad and climbing but still a minority of firms — and concentrated among the largest.

Why the adoption numbers disagree: 18% to 78% depending on who you ask

A Federal Reserve analysis (Allen, FEDS Notes) laid out why headline AI-adoption rates diverge so widely: the Census Bureau's business survey put firm adoption at about 18% at year-end 2025; a real-time population survey found about 41% of workers using generative AI; and a survey of senior leaders estimated 78% of the labor force works at a firm that has adopted AI. The Fed attributed the spread mainly to what each survey samples (firms vs individuals vs labor-force-weighted) and how it frames the question, and flagged a Census methodological change in late 2025 that breaks the earlier trend. The essential caveat for reading every number on this page: adoption is real and rising, but its 'level' depends entirely on the unit of measurement.

First randomized trial of AI in mammography cuts interval cancers 12%

The MASAI trial — the first randomized controlled trial of AI-supported mammography screening, following more than 100,000 women in Sweden — published final results in The Lancet. AI-supported screening detected 81% of cancers at screening versus 74% for standard double-reading, cut the interval-cancer rate (cancers found between screens) by about 12% (1.55 vs 1.76 per 1,000), and left the false-positive rate essentially unchanged, while an earlier analysis found it reduced radiologist reading workload by about 44%. Investigators were deliberately cautious: lead author Kristina Lång (Lund University) called it the first RCT evidence that AI is safe to use in screening, and first author Jessie Gommers (Radboud) noted the results do not yet prove AI reduces the advanced cancers that drive mortality. Peer-reviewed, gold-standard clinical evidence — the counterpoint to the Watson precedent.

S&P Global: AI-project abandonment jumps to 42%

S&P Global Market Intelligence (451 Research) found the share of companies abandoning most of their AI initiatives had risen to 42%, up from 17% a year earlier, and that 46% reported no single enterprise objective had produced strong positive impact — even as about 60% of AI investors had implemented generative AI in production. The figures come from a survey of 1,006 respondents and measure reported behavior and sentiment, not audited outcomes; the most-cited obstacles were data confidentiality, cost and accuracy.

MIT report: 95% of enterprise generative-AI pilots show no measurable return

MIT's NANDA initiative reported in 'The GenAI Divide: State of AI in Business 2025' that about 95% of enterprise generative-AI pilots it studied had produced no measurable profit-and-loss impact, with roughly 5% delivering real value. Tools bought from specialized vendors succeeded about twice as often as those built in-house (about 67% vs 33%), and workforce disruption showed up mainly as unfilled vacancies in customer-support and administrative roles rather than mass layoffs. The finding is a report drawing on 150 leader interviews, a 350-person survey and 300 public deployments — not audited financials — but it is the most-cited reality check on the enterprise-adoption story, and it sits against the U.S. Census Bureau's measured finding that roughly 20% of firms use AI at all: many companies are trying AI; few have yet shown it pays.

Randomized trial: AI slows experienced developers 19% on their own code

In a randomized controlled trial, METR had 16 experienced open-source developers complete 246 real issues on mature codebases they maintain, with and without early-2025 AI tools (Cursor Pro with Claude 3.5/3.7). The developers were 19% slower with AI — even though they had expected a 24% speedup and, afterward, believed AI had sped them up by about 20%. The result sits directly against the 2023 GitHub Copilot trial, which measured a 55.8% speedup: the two randomized experiments reach opposite-signed results because they test different tasks and populations — a narrow greenfield task with mixed-experience recruits versus expert maintainers on complex real code. The gap between measured and perceived productivity is itself the finding.

Klarna rebalances toward humans after an AI-first support push

Klarna began rehiring human customer-service staff after having said AI was doing the work of hundreds of agents, in what several outlets framed as a reversal. Klarna disputed that framing: it says it continues to invest heavily in AI, that its assistant handles the work of more than 800 full-time roles, and that the new hiring is a small quality pilot rather than a retreat. CEO Sebastian Siemiatkowski said an overemphasis on cost-cutting — not AI itself — had hurt service quality. Both the reversal framing and the company's rebuttal are recorded so the contested claim stays auditable.

AlphaFold wins the Nobel Prize in Chemistry

The Royal Swedish Academy of Sciences awarded the 2024 Nobel Prize in Chemistry half to Demis Hassabis and John Jumper of Google DeepMind 'for protein structure prediction' with AlphaFold, and half to David Baker 'for computational protein design.' Committee chair Heiner Linke said the recognized discoveries 'hold enormous potential.' By the announcement, AlphaFold2 had been used by more than two million people in 190 countries. It is the first Nobel Prize to recognize a modern deep-learning system's direct scientific contribution — the strongest possible independent verdict that AI has produced real science.

DeepMind AI reaches silver-medal level at the Math Olympiad

Google DeepMind's AlphaProof and AlphaGeometry 2 together solved four of the six 2024 International Mathematical Olympiad problems for 28 of 42 points — the silver-medal threshold, one point below gold. The solutions were graded by the competition's own markers, including Fields Medalist Timothy Gowers, making it an independently verified demonstration of AI mathematical reasoning rather than a self-reported score.

AlphaFold 3 extends structure prediction beyond proteins

Google DeepMind and Isomorphic Labs published AlphaFold 3 in Nature, a diffusion-based model that predicts the joint structure of proteins together with nucleic acids, small-molecule ligands and ions — the interactions that matter for drug discovery, not just isolated protein shapes. A peer-reviewed step from a scientific result toward a deployable tool.

The U.S. Census Bureau begins measuring business AI use at population scale

The Census Bureau released the first AI supplement to its Business Trends and Outlook Survey, sampling roughly 1.2 million businesses on whether they actually use AI to produce goods and services. It established the country's first population-scale, measured (rather than self-selected or intention-based) record of enterprise AI adoption — the backbone against which vendor and survey claims can be checked.

GNoME predicts 380,000 stable materials — and outside labs make hundreds

In Nature, Google DeepMind's GNoME predicted about 2.2 million new inorganic crystals, roughly 380,000 of them stable enough to be candidates for synthesis. Crucially the result was partly validated outside DeepMind: a companion literature review found external laboratories had already independently synthesized 736 of the predicted structures, and Lawrence Berkeley National Laboratory's autonomous 'A-Lab' made more than 41 using GNoME's stability predictions — independent corroboration of a vendor claim.

GraphCast: AI matches the gold standard in weather forecasting

In a paper in Science, Google DeepMind's GraphCast produced 10-day global weather forecasts in under a minute on a single machine and outperformed the European Centre's gold-standard HRES system on more than 90% of 1,380 verification targets. A peer-reviewed demonstration that AI can rival established physical-simulation methods on an operational scientific task.

The 'jagged frontier': AI helps consultants 40% on some tasks, destroys value on others

A field experiment with more than 750 Boston Consulting Group consultants using GPT-4 (Dell'Acqua, Lakhani et al., later peer-reviewed in Organization Science) found sharply divided results. On creative product-innovation tasks 'inside the frontier,' AI users performed about 40% better, with the biggest gains for the lowest-baseline performers. But on a business-problem-solving task designed to lie just 'outside the frontier' — where the AI was confidently wrong — consultants using GPT-4 were about 23% less likely to reach the correct answer than those without it. The AI's uniform output also cut the diversity of ideas across the group by roughly 41%.

First patient dosed in Phase II of an AI-discovered drug candidate

Insilico Medicine began Phase II trials of INS018_055 (later rentosertib), an idiopathic-pulmonary-fibrosis candidate the company says is the first drug both discovered and designed using generative AI — a claim that is a sponsor characterization, not an independent or peer-reviewed finding, and the drug is not approved. Co-CEO Feng Ren called reaching Phase II a validation of AI-driven drug discovery. Recorded here as a documented milestone in AI's medical pipeline, with its provenance labeled.

The first large workplace study: AI lifts support-agent productivity 14%, most for novices

A field study of 5,179 customer-support agents (Brynjolfsson, Li and Raymond, published through the NBER) found that access to a generative-AI conversation assistant raised issues resolved per hour by an average of 14% — and by about 34% for the least-experienced agents, while barely helping the most experienced. AI assistance also improved customer sentiment, agent retention and how quickly new hires reached proficiency. It is the landmark early evidence that measured workplace value from AI is real but unevenly distributed, concentrated among lower-skilled workers.

Controlled trial: GitHub Copilot users finish a coding task 55.8% faster

In a randomized controlled experiment, software developers given GitHub Copilot completed a self-contained task (writing an HTTP server in JavaScript) 55.8% faster than a control group. The result is one of the earliest clean measurements of AI coding assistance — but on a narrow, greenfield task, a caveat that later real-world studies would sharpen (see the 2025 METR trial).

ChatGPT reaches an estimated 100 million users in two months

Roughly two months after its late-November 2022 launch, ChatGPT was estimated to have reached 100 million users — described at the time as the fastest-growing consumer application in history. The 100-million figure is a third-party estimate (Similarweb, cited by UBS), not an OpenAI-reported metric; it marks the moment consumer AI adoption became a mass phenomenon.

AlphaFold database releases ~200 million predicted protein structures

The AlphaFold team uploaded predicted structures for nearly all catalogued proteins — about 200 million — to the open AlphaFold Protein Structure Database, turning a research result into shared scientific infrastructure. By the time of the 2024 Nobel announcement the database and model had been used by more than two million people from 190 countries, one of the clearest cases of measured real-world scientific adoption.

AlphaFold2 cracks protein-structure prediction at CASP14

DeepMind's AlphaFold2 won the 14th Critical Assessment of Structure Prediction (CASP14) with a median global-distance score around 92.4 out of 100 — accuracy competitive with experimental methods for many proteins. Independent CASP assessors described the roughly 50-year 'protein-folding problem' as essentially solved for a large class of proteins. This is the origin of the modern evidence that AI can produce real, independently graded scientific results.

IBM Watson's cancer-care collapse sets the cautionary precedent

MD Anderson Cancer Center shelved its IBM Watson 'Oncology Expert Advisor' in 2016 after spending a reported $62 million and never deploying it in clinical care. Independent studies later found Watson for Oncology's recommendations matched expert panels unevenly by geography (about 73% concordance in India, 49% in South Korea), and IEEE Spectrum counted nearly 50 Watson Health partnership announcements since 2011, many of which produced no tools in use. It is the reference case for AI hype outrunning delivery in medicine — the failure the rest of this record is measured against.