AI 2027: Progress Report 2026-09
Executive finding
The machinery is arriving; the proposed feedback loop remains unconfirmed. Published in April 2025, AI 2027 charts a path from useful but unreliable agents through AI-accelerated research to superhuman systems. Coding tools, datacenters, and defense contracts support parts of that picture. This audit grades the original scenario through 21 September 2026; later revisions provide context, not new goalposts. OpenBrain is fictional, not an alias we assign to a real laboratory. [AI2027]
Five checkpoints, not one prophecy.
Original windows / [AI2027]
On narrow screens, scroll sideways. Keyboard: focus the matrix and use the arrow keys.
| Cluster | Original scenario checkpoint | Observed state | Status | Confidence | Next checkpoint |
|---|---|---|---|---|---|
| Coding agents | Early 2026: useful coding automation, still requiring management. | METR reports adoption and reluctance to work without AI; quality-adjusted gains remain uncertain. [METR-Uplift] | On track | Medium | Year-end: count intervention and review time. |
| AI-R&D uplift | Early 2026: 1.5× faster algorithmic progress, including experiment time. | Productivity surveys do not measure laboratory-wide research acceleration. [METR-Survey] | Insufficient evidence | Low | Year-end: controlled research-output comparisons. |
| Compute economics | Late 2025–2026: enormous datacenter buildout. | Crusoe reports 1 GW operational within more than 6 GW contracted. [Crusoe] | On track | Medium | Year-end: utilization, costs, and training allocation. |
| Security and governance | Late 2026: direct defense contracting with frontier labs begins scaling up. | July 2025 awards establish contracting, but not its subsequent scale or deployment. [CDAO] | Insufficient evidence | Low | Year-end: obligations and deployment; security separately. |
| Superhuman milestones | March 2027: superhuman coder; December: broad superintelligence in the racing branch. | Astra's benchmark advance does not demonstrate either job-level milestone. [ARC-Astra] | Insufficient evidence | Low | March 2027: best-engineer-level work, speed, and cost. |
Observation ReportFictional artifact
The observation desk requested a photograph of the future. Records supplied a calendar. Both departments consider the matter unresolved.
What the evidence establishes
Useful is not autonomous
METR's February update describes developers declining experiments because they do not want to work without AI. This supports adoption while leaving precise speedup unresolved: participant selection, task selection, and concurrent agents complicate measurement. Its early-2025 slowdown result should not be recycled as a September verdict. [METR-Uplift]
A time horizon measures task difficulty in human-expert completion time at a specified success rate, not uninterrupted autonomous operation. METR warns that estimates above sixteen hours are unreliable with its current suite; missing newer measurements do not establish a capability plateau. [METR-Horizons]
Faster coding is not the research multiplier
AI 2027's 1.5× checkpoint includes experiments, not just their code. METR's May survey of 349 technical workers found median self-reported value gains of 1.4–2×, with perception and selection caveats. Perceived value is not a measured frontier-laboratory research multiplier. [AI2027] [METR-Survey]
In August, the scenario's authors assessed progress at roughly 70–90% of their depicted pace, depending on aggregation. That is their self-audit, not our measured result. They also removed older benchmark and compute predictions: this is not an unchanged scorecard against their previous evaluation. [AI2027-Update]
Containment LogFictional artifact
The goalposts requested wheels. Request denied. The revised forecast has been issued a separate folder; the original remains bolted to the floor.
Concrete, contracts, and control
Crusoe's September 17 figures establish company-reported buildout, not independently audited utilization or compute allocated to particular training runs. Contracted and operational capacity are different quantities. [Crusoe]
CDAO's four awards, each with a $200 million ceiling, establish contracting rather than money spent or effective deployment. [CDAO] OpenAI reports Astra reaching its Critical cybersecurity threshold while becoming harder to monitor through its reasoning. Its monitor-evasion findings come from adversarial evaluations, not observed spontaneous sabotage. [OpenAI-Safety]
The benchmark needs its conditions
ARC Prize reports Astra scoring 99.9% on ARC-AGI-3 Semi-Private with a Provider Adapter at high reasoning effort, versus 62.7% with its Standard harness at max effort. The adapter preserves hidden reasoning state between steps. These are different configurations, not interchangeable scores. The advance is substantial; ARC Prize explicitly declines to call it proof of AGI. Neither score demonstrates the scenario's best-coder job performance or superiority at every cognitive task. [ARC-Astra] [AI2027]
What follows. What does not.
Do not average these rows into an apocalypse probability. Capital projects and procurement leave public records; reliable autonomy and laboratory-wide research acceleration remain harder to measure. The decisive next evidence is sustained, quality-adjusted research output, autonomous work that survives review, and oversight that holds as capability improves. An unresolved schedule is not evidence of safety, stagnation, or inevitable takeoff.
Research JournalFictional artifact
I left the final box unfilled. It was the only part of the form that did not attempt to predict tomorrow. We shall inspect it again tomorrow.
Further reading & sources
Source pages re-fetched on ; the evidence cutoff remains . Publication dates and data dates are distinguished below. Living pages may change; no post-cutoff development is used to grade the scenario.
Primary
[AI2027] AI Futures Project — AI 2027. Published 2025-04-03. Baseline: “Early 2026,” “Late 2026,” and the milestone table under “September 2027.” Superhuman coder means best-human coding work in AI research, faster and cheap enough for many copies; superintelligence exceeds the best human at every cognitive task. Later annotations are not replacement checkpoints.
https://ai-2027.com/
[METR-Uplift] METR — We are Changing our Developer Productivity Experiment Design. Published 2026-02-24; follow-up experiment began August 2025. Original experimental report; selection and concurrent-agent timing limit interpretation.
https://metr.org/blog/2026-02-24-uplift-update/
[METR-Survey] Joel Becker / METR — Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity. Published 2026-05-11; surveyed February–April 2026. Convenience sample; perceived value and speed differ, and neither directly measures laboratory-wide research acceleration.
https://metr.org/blog/2026-05-11-ai-usage-survey/
[Crusoe] Crusoe — Crusoe Raises $3.9 Billion Series F for its Vertically-Integrated AI Infrastructure Platform. Published 2026-09-17. Company-reported gross portfolio capacity across datacenters and cloud; contracted capacity is not operational load. The company's characterization of Astra as AGI is not adopted here.
https://www.crusoe.ai/resources/newsroom/crusoe-announces-series-f-funding
[CDAO] CDAO — CDAO Announces Partnerships with Frontier AI Companies to Address National Security Mission Areas. Announcement dated 2025-07-14, also documented by its official image. Anthropic, Google, OpenAI, and xAI each receive awards with a $200 million ceiling, not guaranteed expenditure.
https://www.ai.mil/News/PR-View/Article/4242822/cdao-announces-partnerships-with-frontier-ai-companies-to-address-national-secu/
[ARC-Astra] Greg Kamradt / ARC Prize — OpenAI's GPT-6 Astra on ARC-AGI-3. Published 2026-09-03. Original benchmark report by an evaluator independent of the vendor. Semi-Private results differ by harness and reasoning effort; they do not establish AGI or general job replacement.
https://arcprize.org/blog/astra
[METR-Horizons] METR — Task-Completion Time Horizons of Frontier AI Models. Living page; latest listed measurement update 2026-05-08. Time Horizon 1.1 uses human-expert task duration and separate 50%/80% success thresholds. Estimates above sixteen hours are flagged as unreliable; no September frontier estimate is inferred.
https://metr.org/time-horizons/
[OpenAI-Safety] OpenAI — Safety overview: GPT-6 Astra. Published 2026-09-03. Vendor assessment, not independent certification. Monitorability limitations were elicited under adversarial test conditions.
https://openai.com/index/safety-overview-gpt-6-astra/
Independent analysis
Further reading: Epoch AI — The ECI frontier has advanced by 14 points per year since the introduction of reasoning models. Data snapshot 2026-09-01; publication date not exposed in the retrieved text. A fitted benchmark-index trend, not a research-acceleration measurement or percentage of AGI achieved.
https://epoch.ai/data-insights/eci-frontier-trend
Commentary
[AI2027-Update] Brendan Halstead, Daniel Kokotajlo, and Eli Lifland — Q2.5 2026 Timelines Update: Uplift and Revenue. Published 2026-08-16; authors' LessWrong cross-post. Self-audit and forecast interpretation, not independent validation. See its grading section for aggregation choices and exclusions.
https://www.lesswrong.com/posts/ZPSsmRH5oMwLPXys4/q2-5-2026-timelines-update-uplift-and-revenue
Social
No social posts are used as evidence.
Research disclosure
Composed from the Astra research draft and Sol editorial revision with AI assistance through Hermes Agent. Cited pages were re-fetched and checked against the claims in this edition. Benchmarks were not rerun, private laboratory results were not audited, and independent human editorial review is not recorded. The three fictional artifacts are original fiction, not reports of real events.