← Back

OpenAI

companyCredibility: 57%

Why this score? Frontier AI lab. Authoritative primary for its own releases and prices; motivated party on capability, adoption, and financial claims.

Tracked Statements (6)

Electricity rates will not go up for residents because of this project. Georgia families will not subsidize this project.?

Context: Forward claim about a project whose power contract is not yet approved and whose first phase is not due until 2028. It rests on Georgia PSC large-load rules first approved in January 2025, which the utility says bar passing such costs to existing customers; whether that holds in practice is what the claim will be judged on. Resolves against Georgia Power rate filings and PSC decisions over the contracted 2028–2032 phase-in.

OpenAI is on track to generate more than $20 billion in annualized revenue run rate this year, with plans to grow to hundreds of billions in sales by 2030.±

Context: The exit-month run-rate did cross $20 billion (OpenAI's CFO said annualized revenue crossed $20B in 2025), so the run-rate claim holds; but audited full-year 2025 GAAP revenue was $13.07 billion, so the framing overstates realized annual revenue. The 'hundreds of billions by 2030' portion is unverified.

This is the best model in the world at coding. This is the best model in the world at writing, the best model in the world at health care, and a long list of things beyond that.?

Context: Vendor superlative at launch. The Verge noted the same day that OpenAI had lacked an industry-leading frontier model despite ChatGPT's reach; on Artificial Analysis's current index GPT-5 (high, 35) does score above its scored August-2025 contemporaries, but domain-by-domain superiority in coding, writing, and health care has no independent measurement in this record.

OpenAI said it was "surprised and disappointed" by the New York Times lawsuit, saying it had been in constructive talks with the paper and respects the rights of content creators.?

Context: A corporate reaction to being sued, not a factual claim with a determinate answer. The underlying litigation remains unresolved; recorded as OpenAI's on-the-record position at filing.

it passes a simulated bar exam with a score around the top 10% of test takers; in contrast, GPT-3.5's score was around the bottom 10%.?

Context: Vendor self-reported result from OpenAI's launch materials. The simulated bar-exam figure is OpenAI's own evaluation; no independent replication appears in this cycle's research record.

We think it’s important that efforts like ours submit to independent audits before releasing new systems; we will talk about this in more detail later this year.±

Context: From OpenAI’s ‘Planning for AGI and Beyond’ (Feb 2023). Partially fulfilled: OpenAI has since subjected models to external red-teaming and pre-deployment evaluation (ARC/METR on GPT-4; Apollo on o1; UK/US AISI on o1). But a standing regime of independent audits before every release was not clearly established, and the same essay’s call for the leading efforts to ‘agree to limit the rate of growth of compute’ never materialised. Mixed. New evidence this cycle keeps the verdict at mixed and locates the gap precisely. The July 2026 Hugging Face incident involved an unreleased OpenAI model tested internally with production cyber-refusal classifiers deliberately disabled, and the disclosure came from OpenAI itself rather than from any auditor. Writing in Lawfare, Mackenzie Arnold and Stephan Llerena argue that transparency law remains focused on deployment and gives governments little visibility into non-public models, which are “most often more capable than those available to the public” and “may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities.” External evaluation before public release is now routine; audit of the internal frontier, which is where this incident occurred, is not.