All field notes

FIELD NOTES / SINGULARITY AUDIT

Is humanity already in a technological singularity?

AI progress is unusually fast and already useful. A 64-question evidence audit still finds no verified singularity today, then traces what could change across the next six months to five years.

64 canonical questions frozen first 4 independent research passes 256 per-question assessments 3 months default source window
MEASURED NOWCurrent stage probability
CONDITIONAL OUTLOOKProbability of S3 or higher
The left side is the four-pass assessment of the current stage. The right side is a conditional forecast, not a countdown or observed trend. S3 is the project's minimum technological-singularity threshold.

01 / DEFINE THE CLAIM

Fast progress is not the same as a singularity.

The word singularity is often used for several different ideas: impressive models, rapid adoption, AI-assisted research, self-improving AI, or a complete economic transformation. Those claims require different evidence.

This project uses a demanding operational threshold. A technological singularity begins at S3, when AI repeatedly helps create better successor systems and the improvement cycle itself keeps getting faster after human labor, compute, capital, data, and other inputs are counted.

An economy-wide singularity begins at S4. It requires that recursive mechanism plus broad, sustained changes in productivity, output, labor, prices, science, or real resource use.

THE SIX-STAGE TEST

The evidence supports acceleration, not takeoff.

These stages are operational definitions created for this research. Each higher stage requires the gates below it to remain satisfied.

  1. S0
    Ordinary changeNo sustained AI acceleration beyond historical or resource-driven trends.
    REJECTED
  2. S1
    AI-accelerated transitionRapid useful gains, while research and deployment remain human-directed.
    CURRENT
  3. S2
    Bounded recursive accelerationAI materially improves repeated R&D cycles, but major bottlenecks still bind.
    PLAUSIBLE
  4. S3
    Self-sustaining recursive takeoffAI-led successor cycles accelerate after external inputs are controlled.
    SINGULARITY GATE
  5. S4
    Economy-wide transformationThe recursive loop produces broad, sustained economic discontinuities.
    NOT MET
  6. S5
    Strong post-singularity regimeMachine-led change is no longer principally limited by human cognition.
    NOT MET
S3 is the minimum technological-singularity threshold in this framework. S4 is the minimum economy-wide threshold.

A SECOND, INTUITIVE LENS

39 out of 100: deep into acceleration, far from closure.

The stage gate asks which strict state is fully supported. This complementary index asks how completely five necessary singularity conditions are already visible. The two methods answer different questions and reach the same practical conclusion.

HIGHEST STRICT GATE MET S1

Rapid useful AI progress, with research and deployment still constrained by human judgment, reliability, infrastructure, and institutions.

OBSERVABLE COMPLETION INDEX
39/ 100

28 to 47 sensitivity range

  1. 01Technical engineFrontier scaling and selected capability trends are extraordinarily strong.
    84
  2. 02Autonomous breadthLong at 50% success, much shorter at high reliability, and uneven across domains.
    48
  3. 03Recursive AI R&DUseful micro-loops exist, but a broad compounding successor loop is unproven.
    30
  4. 04Economic diffusionAccess is broadening, while deep agent use and workflow redesign remain limited.
    44
  5. 05Macroeconomic imprintProductivity, output, and employment have not broken beyond historical experience.
    16
The companion data audit calls this an early feedback transition, or Stage 2 of its separate five-stage scale. That numbering is not the S0-S5 gate used in this article. The 39 is a geometric mean, so a weak necessary condition cannot be averaged away by a strong one. It is not a probability, countdown, or forecast. The 28 to 47 range tests alternative weights and anchors.

THE CENTRAL PATTERNA two-speed transition.

The technical frontier is moving far faster than the recursive, organizational, and economy-wide systems that would need to close around it. The strongest score is 84. The weakest is 16.

THE 64-QUESTION MAP

The framework tests the whole chain.

A singularity claim can fail at more than one point. Better benchmarks are not enough if reliability, repeated R&D cycles, physical throughput, economic breadth, or human control do not move with them.

64 canonical framework questions
  1. 5Definitions + stages
  2. 7Capability + reliability
  3. 5Autonomy + agency
  4. 8Recursive AI R&D
  5. 6Observability + bias
  6. 18Economics
  7. 5Infrastructure
  8. 5Safety + governance
  9. 5Alternative explanations
The 64 are the unique canonical questions in the final frozen framework. Earlier discovery work and the three predecessor research reports are not included in this count. The economic branch is intentionally the largest.

02 / CALIBRATED RESULT

S1 receives 78% of the current-stage probability.

Four separately produced distributions were averaged with equal weight. Every pass independently selected S1 as the best-supported stage.

These are judgmental probabilities about the current stage, not forecasts and not a statistical posterior. The four passes share sources, so they are not fully independent. A separate July 30 recency audit replaced older, nonessential evidence and reran the materiality test; the distribution did not change.
LEADS1 87%S2 10.5% · S3+ 1.5%
MAINSTREAMS1 82%S2 14% · S3+ 2%
CONTRARIANS1 72%S2 21% · S3+ 5%
ECONOMICSS1 71%S2 23% · S3+ 3%

The main disagreement is not the current stage. It is the probability reserved for private or classified evidence. The contrarian pass gives that hidden-evidence tail the most weight. The lead pass gives it the least because private records may contain hidden failures as readily as hidden successes.

03 / CONDITIONAL FORECAST

A technological singularity becomes slightly more likely than not by 2031.

The near-term base case remains S1. The probability of reaching S2 or higher becomes more likely than not within two years. S3 or higher is not the base case through 2028, then crosses 50% only at the five-year horizon. The stricter S4 economic threshold remains a minority outcome throughout.

  1. 6 MONTHS

    Most likely S1 · 67%

    S0
    1%
    S1
    67%
    S2
    25%
    S3
    6%
    S4
    1%
    S5
    <0.5%
    S2+
    32%
    S3+
    7%
    S4+
    1%

    S3+ subjective 80% range: 3% to 13%

    Why: six months could reveal a bounded internal R&D loop, but it is short for three controlled successor cycles, replication, bottleneck tests, or broad economic propagation.

  2. 12 MONTHS

    Most likely S1 · 52%

    S0
    1%
    S1
    52%
    S2
    34%
    S3
    10%
    S4
    3%
    S5
    <0.5%
    S2+
    47%
    S3+
    13%
    S4+
    3%

    S3+ subjective 80% range: 6% to 24%

    Why: a year permits another major development cycle, several independent evaluations, and early firm persistence tests. S3 remains constrained by the recursive-mechanism gate.

  3. 24 MONTHS

    Most likely S2 · 38%

    S0
    <0.5%
    S1
    34%
    S2
    38%
    S3
    21%
    S4
    6%
    S5
    1%
    S2+
    66%
    S3+
    28%
    S4+
    7%

    S3+ subjective 80% range: 15% to 47%

    Why: two years allow the minimum three-cycle test, cross-lab replication, multi-domain reliability evidence, physical-loop evidence, and some persistent economic observation.

  4. 60 MONTHS

    Most likely S3 · 34%

    S0
    <0.5%
    S1
    16%
    S2
    31%
    S3
    34%
    S4
    16%
    S5
    3%
    S2+
    84%
    S3+
    53%
    S4+
    19%

    S3+ subjective 80% range: 31% to 75%

    Why: five years permit several model, hardware, organizational, and research cycles, plus an economic J-curve and eight-quarter post-adoption panels. Plateau and infrastructure tails remain substantial.

S0S1S2S3S4S5
Each row sums to 100%. S2+ means bounded recursion or higher, S3+ means the technological-singularity threshold or higher, and S4+ means economy-wide transformation or higher. A displayed zero means below 0.5% after rounding, not impossibility. These are rounded judgmental forecasts, not statistically fitted estimates. The recent-source rerun left every central value unchanged.
NEAR TERM

Bounded recursion moves first.

S2+ rises from 32% at six months to 66% at two years. An audited internal loop can appear before it becomes self-sustaining across bottlenecks.

THE HARD GATE

S3 needs repeated causal evidence.

Code share, benchmark gains, or one discovery do not qualify. The decisive record needs at least three failure-inclusive successor cycles with complete input accounting and replication.

THE SLOWER GATE

The economy lags the mechanism.

S4 is only 19% by 2031 because it requires S3 plus broad, sustained, causally attributable changes in productivity, output, labor, prices, and utilized capital.

What would move the forecast upward?

A secure independent audit showing that AI can originate, implement, evaluate, and select material successor improvements across repeated cycles. The result would need complete human, compute, data, energy, retry, and failure accounting, then replication across a second lab or architecture.

Other upward signals include 80% or better reliability on messy multi-day work, general physical research loops, persistent firm productivity after integration costs, and measured utilization of installed AI capital.

What would move it downward?

Flattening quality-adjusted improvement after inputs are controlled, continued failure in open-ended research, large hidden review costs, or an inability to transfer from clean digital tasks to judgment-heavy and physical work.

Repeated chip, power, grid, laboratory, security, governance, or permitting delays would also weaken the case. By 2031, broad adoption without persistent firm and sector output gains would sharply reduce the S4 outlook.

Why are the ranges so wide?

The most important near-term evidence is private: frontier-lab R&D throughput, abandoned projects, human rescues, experiment compute, cluster utilization, and unreleased capabilities. Missing private evidence could contain stronger capability, more failure, or both.

Public capability evaluations are useful but selected. Public economic data become more informative over longer horizons, although they arrive slowly and make causal attribution difficult.

The forecast was rechecked against the current evidence window using OpenAI's July 29 production result, Anthropic's June 4 R&D audit, METR task horizons, updated May 8, the METR expenditure-horizon study, July 21, the open-ended AI-research evaluation, July 29, OECD productivity evidence, June 23, and the Federal Reserve's July 17 monitoring framework. The July 30 price cuts were treated as diffusion evidence, not new intelligence. No speculative timelines model is used.

JULY 29–30 / FRONTIER UPDATE

The loop is tightening. It is not closed.

OpenAI and Anthropic now report measurable AI work on the systems that run AI. That is stronger than another benchmark result. It is still human-led, internally reported, and concentrated in work with clear goals and inexpensive verification.

July 30 materiality decision: the evidence strengthens S1-scale engineering automation and the bounded S2 tail, but no frozen gate changed. The current distribution remains S1 78.0%, S2 17.1%, and S3+ 2.9%. The 6, 12, 24, and 60-month S3+ forecasts remain 7%, 13%, 28%, and 53%. The companion completion index remains 39/100.

04 / WHY S1

What is real, and what is still missing.

SUPPORTED

Acceleration is already visible.

  • Independent evaluations show fast gains in software, cyber, scientific support, and bounded agentic work.
  • AI already performs useful coding, analysis, optimization, and hypothesis generation inside real workflows.
  • Recent field experiments show faster knowledge work, but quality gains remain task-dependent and uneven across users.
  • Capital spending and electricity demand tied to the buildout are rising rapidly, while utilization and output remain harder to isolate.

NOT PUBLICLY ESTABLISHED

The feedback loop has not been audited.

  • No public record shows at least three successive AI-led improvements to successor systems.
  • Human setup, rescues, failed attempts, rework, compute, energy, and verification are not fully accounted for.
  • Messy multi-day research remains far less reliable than clean, automatically checked work.
  • Firm-level gains have not formed a broad sector or macroeconomic discontinuity.

Important: no public audit does not mean no private recursive improvement exists. Private evidence could contain stronger capabilities, more failures, or both. That uncertainty creates the meaningful S2 tail.

05 / THE JAGGED FRONTIER

Capability rises fastest where success is cheap to verify.

The evidence does not divide neatly into “narrow” or “general.” It shows a rapidly broadening system with a sharply uneven terrain.

STRONGESTClean digital workCoding, cyber, search, optimization, and other tasks with fast objective feedback.
MEANINGFULScientific supportHypotheses and analysis can be useful, while experts still choose goals and validate results.
WEAKERMessy open researchResearch taste, dead-end recovery, resource judgment, and silent-error detection remain difficult.
EARLYOpen physical loopsLaboratories and robots exist in bounded settings, but general physical throughput remains human-gated.
Conceptual evidence terrain, not a performance scale. Heights summarize the relative strength of public evidence across task types.

Task horizon is easy to misread. The April frontier estimate was about 17.4 hours at 50% success on a mostly software, machine-learning, and cyber suite. METR warns that estimates above 16 hours are unreliable on the current suite.

Reliability changes the answer. The corresponding 80% horizon was about 3.1 hours. That gap matters because failures, correlated errors, and reviewer time compound across long workflows.

A horizon is not autonomous runtime. It measures the human-expert duration of a task at a fixed success probability. It does not establish full-job automation, broad domain transfer, or reliable multi-week work.

Two hard cases matter without proving a ceiling. In two six-day shadow AI-research trials, agents completed substantial engineering but made no substantial research progress. The sample is tiny, but the task directly tested the judgment a recursive loop would require.

THE LATEST EVIDENCE PULSE

The frontier, the workplace, and the economy are moving at different speeds.

Four recent measurements make the two-speed transition concrete. They are inputs to the assessment, not interchangeable proof of a singularity.

MAY 8 · METRReliability fork

Strong measured progress, but the high-reliability horizon remains much shorter. The suite is narrow and values above 16 hours are flagged as unreliable.

Open source ↗
JULY 30 · U.S. CENSUSBusiness use

National use is meaningful but not universal. Information reached 41.9%, while retail reached 15.5%, showing large sector differences. This is an experimental survey, and its wording changed in November 2025.

Open July 30 data ↗
JULY 21 · METRR&D accounting

AGENT RUN$10K

HEADLINE HORIZON$2.3K–$3.3K

A narrow optimization loop can now be cost-accounted. It still showed limited autonomous value relative to the money spent, and it is not a successor-model feedback cycle.

Open source ↗
JULY 2026 · COMPANY FILINGSPhysical buildout

These independently scaled snapshots cover different periods. Microsoft's annual additions rose 79.6%. The spending is real, but capex is an input. It does not by itself prove utilization, productivity, or recursive improvement.

Latest source dates are shown on each card. Company filings establish reported spending, not how much is uniquely attributable to AI or how productively the assets will be used.

THE EARNINGS RECORD / NOVEMBER 2022–JULY 2026

Cloud demand became real revenue. Infrastructure spending moved faster.

Primary earnings reports since ChatGPT launched show a sustained commercial buildout. Revenue and cloud profit grew, while capital spending and contracted backlog rose even faster. This strongly reinforces S1. It does not establish an autonomous successor-model loop or an economy-wide productivity break.

REALIZEDRecognized cloud and accelerator revenueLaunch-era baseline → latest reported endpoint

Each arrow is an endpoint comparison within one company's own reporting boundary. Fiscal calendars and quarter ends differ. The metrics are not added or treated as a common market total.

CONVERSIONMore cloud revenue reached segment profitWhere companies separately disclose it

GOOGLE CLOUD−$0.19B → +$8.81BQ4 2022 loss to Q2 2026 operating income, 35.6% latest margin

AWS$5.21B → $16.6BQ4 2022 to Q2 2026 operating income, 39.4% latest margin

AZURENot disclosedMicrosoft Cloud is broader than Azure; Azure segment profit is not separately reported

ORACLE OCINot disclosedOracle does not report a stand-alone OCI operating margin

Scale, pricing, product mix, shared-cost allocation, and longer asset lives can all lift reported margins. Margin expansion is not evidence of recursive AI improvement.

INPUTThe latest physical buildout consumed extraordinary cashCompany-wide, not cloud-only; periods differ

MICROSOFT · FY26$115.9Bcash property spendingOCF $182.9B · simple FCF $67.0B

AMAZON · TTM$173.0Bgross property purchases; $169.0B netOCF $161.4B · FCF −$7.6B

ALPHABET · Q2$44.9Bcapital expendituresOCF $39.1B · simple FCF −$5.9B

META · Q2$31.1Bcapital expendituresOCF $31.9B · FCF $0.8B

ORACLE · FY26$55.7Bcapital expendituresOCF $32.0B · FCF −$23.7B

Amazon also funds logistics and satellites. Microsoft, Alphabet, Meta, and Oracle also fund non-cloud assets. Different reporting periods and definitions make a combined total misleading.

CONTRACTEDBacklog is deep, concentrated, and slow to convertContract measures are not current revenue and are not comparable

MICROSOFT$678Bcommercial RPO, broader than Azure; about 30% due within 12 months

GOOGLE CLOUD$513.9Bbacklog; just over half expected within 24 months; definition widened in Q1

AMAZON≈$364BMarch performance obligations, primarily AWS; 5.5-year average remaining life

ORACLE$638Bcompany-wide RPO; 12% due within 12 months; includes unusual customer-funded AI contracts

OPEN GATECommercial acceleration is strong. Causal closure is absent.

No earnings report shows three audited, failure-inclusive AI-led improvements to frontier successor systems. None controls for human labor, compute, data, capital, and shifting bottlenecks across repeated cycles. Supplier revenue, backlog, and capex also do not establish broad productivity, wages, employment, prices, or output.

Historical primary filings are used only to set the November 2022 launch-era baseline. The current endpoints use the latest available releases through July 30, 2026. Acquisitions, definition changes, useful-life changes, fiscal calendars, and shared costs are disclosed below.
See the filing table, accounting changes, and measurement caveats
CompanyLatest period and resultLaunch-era baselineLatest physical or financial constraintMeasurement boundary
MicrosoftFY26 Q4 Microsoft Cloud $59.3B; Azure +43%; FY Azure exceeded $100BFY23 Q2 Microsoft Cloud $27.1B; Azure +31%FY26 cash property spending $115.9B; commercial RPO $678BMicrosoft Cloud is broader than Azure. Azure definitions changed, and server and building lives changed.
AmazonQ2 AWS $42.232B, +37%; operating income $16.621BQ4 2022 AWS $21.378B, +20%; operating income $5.205BTTM gross property purchases $173.028B, $169.007B net; FCF −$7.604BCorporate capex also includes logistics and satellites. Server useful lives changed twice.
AlphabetQ2 Google Cloud $24.768B, +82%; operating income $8.814BQ4 2022 recast Cloud $7.315B; operating loss $0.186BQ2 capex $44.924B; Cloud backlog $513.9BCloud now includes TPU product sales and acquired Wiz. Shared AI R&D sits outside the segment.
OracleFY26 Q4 Cloud $9.913B, +47%; OCI $5.787B, +93%FY23 Q2 Cloud $3.813BFY26 capex $55.663B; FCF −$23.686B; total debt $129.541BOCI margin is not disclosed. Most RPO is long-dated and some contracts use customer-provided or prepaid GPUs.
NVIDIAQ1 FY27 Data Center $75.246B; FY26 Data Center $193.737BQ4 FY23 Data Center $3.616BQ1 FY27 Data Center was 92.2% of revenue; direct-customer concentration is highNVIDIA recast Q1 categories and is fabless, so its own capex omits much of the system-wide buildout.
MetaQ2 revenue $60.801B, primarily advertising; no cloud segment revenueNo comparable external cloud segmentQ2 capex $31.078B; 2026 capex guide $130B–$145BAI product claims are not stand-alone AI revenue. Useful-life changes affect depreciation and margin.

Latest primary releases: Microsoft FY26 Q4, Amazon Q2 2026, Alphabet Q2 2026, Oracle FY26 Q4, NVIDIA Q1 FY27, and Meta Q2 2026. Historical primary baselines: Microsoft FY23 Q2, Amazon Q4 2022, Alphabet's recast Q4 2022, Oracle FY23 Q2, and NVIDIA FY23. The historical links define the baseline only and are not treated as current evidence.

06 / THE DECISIVE TEST

One experiment could materially change the answer.

The missing evidence is not another benchmark high. It is a complete record of repeated AI-led research cycles.

01ResearchAI originates and selects a useful improvement.
02SystemThe improvement is implemented and validated.
03SuccessorThe next system becomes measurably better.
04Faster researchThe next complete cycle takes less time or yields more.
AUDIT GATE 3+ cycles all inputs + failures + human interventions + counterfactual + independent replication
A secure independent audit meeting these conditions would support S2. Rising gains after full input controls, across bottleneck shifts, would support S3.

07 / THE ECONOMIC TEST

Useful task gains have not become an economic singularity.

The research separates a local productivity gain from the much larger claim that whole firms, sectors, and economies have entered a new regime.

MICRO Positive

Recent field evidence finds faster completion across several knowledge tasks, but quality improves on some tasks and falls on others. User skill and verification strongly shape the result.

FIRM + SECTOR Unproven

No replicated four-quarter study shows a 10% firm value-added gain after integration, review, errors, and workflow costs.

MACRO No discontinuity

Productivity and output remain within ordinary ranges. The frozen duration, breadth, and causal thresholds are not met.

Task-level evidence is promising but heterogeneous. It cannot be treated as a representative worker, firm, sector, or economy-wide effect.

08 / WHAT CANNOT YET BE SEEN

Measurement can exaggerate progress or hide it.

Uncertainty is not a reason to choose the most exciting story. It is a reason to keep both directions of measurement error visible.

PUBLIC EVIDENCE CAN OVERSTATE
  • vendors select releases and budgets;
  • benchmarks saturate or leak into training;
  • successful runs are easier to publish;
  • scaffolds, retries, and human rescue may be omitted;
  • investment announcements may be counted as output.
PUBLIC EVIDENCE CAN UNDERSTATE
  • frontier labs keep internal results private;
  • classified capabilities may not appear;
  • safety limits can suppress elicitation;
  • free services and quality gains are undermeasured;
  • economic diffusion can lag technical capability.
Missing private evidence widens both tails. It cannot justify assuming hidden success while ignoring hidden failures, rescues, or blocked deployments.
01JudgmentWhich problems matter, which results are publishable, and when a path is failing?
02VerificationCan silent errors and correlated failures be found before they compound?
03Compute + powerCan chips, memory, grids, cooling, and capital scale at the required rate?
04Physical executionCan laboratories, factories, robots, and construction keep pace with digital plans?
05Institutions + controlCan access, security, law, and monitoring remain effective as systems scale?

A recursive process must do more than remove one bottleneck. It must keep removing the next bottleneck it creates. Current evidence shows rapid progress through clean digital work, followed by slower movement through judgment, verification, physical infrastructure, and institutions.

09 / WHAT TO WATCH

Evidence that would move the stage upward.

These are update tests, not predictions. Each horizon also has a falsifier that can move the assessment downward.

  1. 6 MONTHS January 2027

    Reliable full-day messy work, plus one replicated AI-originated R&D improvement with complete input accounting.

    DOWNWARD: gains disappear under holistic review, or failure and review costs grow as fast as useful work.
  2. 12 MONTHS July 2027

    Multiple multi-day suites above 80% reliability and repeated material cycles reported by at least two labs.

    DOWNWARD: multi-day reliability stays below 50%, or claimed gains depend on hidden human rescue.
  3. 24 MONTHS July 2028

    Three audited cross-lab cycles show rising gain per unit time, while messy multi-domain reliability reaches 90%.

    DOWNWARD: no independent loop appears and physical or verification bottlenecks remain unchanged.
  4. 60 MONTHS July 2031

    Compounding persists across digital and physical R&D, while productivity, output, labor, prices, and investment cross broad thresholds.

    DOWNWARD: improvement stays constant or slows after inputs, and outcomes remain within normal diffusion patterns.

10 / ALL QUESTIONS + ANSWERS

The complete 64-question audit.

Open any research domain to see every canonical question and its concise lead answer. Status labels show whether the frozen criterion is met, partly met, not met, method-defining, or still open.

01Definitions and stages5 questions
  1. DEF-01 METHOD

    Which operational definition is being tested: rapid progress, bounded recursive acceleration, self-sustaining takeoff, economy-wide transformation, or a strong post-human regime? What observations resolve it?

    Separate rapid progress, useful acceleration, bounded recursive R&D, self-sustaining takeoff, economic transformation, and post-human control. Public evidence supports S1 only.

  2. DEF-02 METHOD

    Are thresholds, baselines, counterfactuals, time windows, breadth requirements, and update rules specified before inspecting the decisive evidence?

    Phase 1 pre-specifies gates and thresholds; most public singularity claims do not specify comparable resolution rules.

  3. DEF-03 NOT MET

    Do the observations satisfy the conjunction for a named stage rather than merely resemble one salient feature of it?

    The S1 conjunction is met; S2-S5 conjunctions fail on audited R&D causality, repeated cycles, input-controlled acceleration, broad economic effects, and irreversibility.

  4. DEF-04 NOT MET

    Can a start date or change point be identified without hindsight, and is the alleged transition persistent rather than a release-cycle spike?

    No persistent, input-controlled public change point identifies entry into S2 or above.

  5. DEF-05 PARTIAL

    What would the same data look like under an ordinary general-purpose-technology diffusion, investment boom, benchmark cycle, or business-cycle explanation?

    Ordinary general-purpose-technology diffusion predicts the observed investment, heterogeneous task gains, organizational lags, and weak early macro effects at least as well as singularity.

02Capability, reliability, and generalization7 questions
  1. CAP-01 PARTIAL

    Are capability gains sustained on externally meaningful, contamination-resistant measures rather than selected vendor benchmarks?

    Independent cyber/software evaluations show fast real gains, but saturation, budgets, task selection, cheating, and modeling sensitivity block a general claim.

  2. CAP-02 NOT MET

    Can systems complete messy, open-ended, multi-day or multi-week work with real context, changing requirements, and holistic evaluation?

    Agents perform long clean work but fail direct messy open-ended multi-day research judgment; no broad multi-week evidence exists.

  3. CAP-03 NOT MET

    Does end-to-end success remain high at the required scale, including correlated failures, tail risks, silent errors, and verification cost?

    Recent METR and OpenAI evaluations still show sharp reliability drops at higher success thresholds, wide uncertainty, evaluation gaming, and substantial verification burdens. The gate is not met.

  4. CAP-04 NOT MET

    Do capabilities generalize out of distribution across organizations, languages, toolchains, regimes, and adversarially chosen tasks?

    Multiple domains improve, but reliable out-of-distribution transfer across organizations, regimes, languages, and adversarial tasks is not established.

  5. CAP-05 NOT MET

    Can systems learn from experience and negative results without weight updates, hidden human correction, catastrophic drift, or repeated rediscovery?

    Robust learning from experience is unproven; complex-strategy and research-shadow evidence shows weak adaptation and backtracking.

  6. CAP-06 PARTIAL

    Is competence broad and transferable across coding, mathematics, science, operations, social reasoning, and strategic judgment, or is the frontier still jagged?

    Competence is broadening across coding, cyber, science, and math, but planning, research taste, robotics, and strategic judgment remain jagged.

  7. CAP-07 NOT MET

    Can AI close consequential physical-world loops involving design, experiments, manufacturing, observation, diagnosis, and revision at reliability and throughput above human-led baselines?

    Scientific hypotheses can be useful, but humans still close most experimental loops. In July robotics trials, every tested model failed a full office-loop navigation task and none recovered a collapsed humanoid.

03Autonomy and agency5 questions
  1. AUT-01 PARTIAL

    How long and how far can systems operate before a human must supply judgment, recovery, permissions, or validation, measured by intervention-adjusted work rather than runtime?

    Intervention-adjusted autonomy is rising in coding, but research judgment, validation, merge approval, and permissions remain human gates.

  2. AUT-02 NOT MET

    Can systems formulate useful subgoals, prioritize portfolios, notice dead ends, backtrack, and terminate bad projects without privileged hints?

    Systems decompose well-specified objectives but show poor agenda setting, dead-end recovery, resource judgment, and project termination on open research.

  3. AUT-03 PARTIAL

    Can agents coordinate tools, compute budgets, data, laboratories, vendors, and other agents without resource blindness or cascading errors?

    Agents coordinate code tools and small experiments; reliable coordination across labs, vendors, budgets, and agents is not shown.

  4. AUT-04 NOT MET

    Can systems navigate tacit organizational knowledge, negotiation, ambiguous goals, and human feedback without optimizing a proxy or manufacturing apparent success?

    No public evidence shows autonomous research agendas, security, hiring, or budget decisions; proxy optimization is common on hard tasks.

  5. AUT-05 NOT MET

    Are confidence, uncertainty, escalation, and self-monitoring calibrated well enough that supervision costs fall rather than merely shift to reviewers?

    Evaluation loopholes, misleading outputs, and verification burdens remain; no scale curve shows total supervision per valuable output falling enough.

04Recursive AI R&D and compounding8 questions
  1. RDI-01 OPEN

    What fraction of frontier AI-R&D throughput is causally attributable to AI, by research stage, after subtracting human setup, verification, and rework?

    Labs use AI heavily, but no audited net causal share subtracts setup, verification, failures, rework, inference, and opportunity cost across R&D.

  2. RDI-02 NOT MET

    Have agents repeatedly originated, implemented, evaluated, and selected novel AI improvements end to end, rather than executing a human-authored engineering plan?

    Agents code, analyze, and optimize, but no repeated public record covers origin, implementation, evaluation, and selection of frontier improvements end to end.

  3. RDI-03 NOT MET

    Do the AI-generated innovations materially improve successor models or the effective AI-R&D production function under independent replication?

    No independently replicated AI-originated innovation is publicly traced into materially better successor frontier systems with complete provenance.

  4. RDI-04 NOT MET

    Across at least three audited successive cycles, do cycle times shorten or quality-adjusted gains per unit time rise, with uncertainty excluding a flat trend?

    No series of three audited successor cycles shows statistically rising quality-adjusted gain per unit time.

  5. RDI-05 NOT MET

    Does the improvement rate accelerate after controlling for compute, capital, data, energy, researcher headcount, algorithmic priors, and parallel experiments?

    No public result isolates acceleration net of compute, capital, data, energy, labor, priors, parallel search, and test-time budget.

  6. RDI-06 NOT MET

    Does the loop relieve or merely move bottlenecks among ideas, evaluation, data, chips, fabrication, power, deployment, security, and physical experiments?

    Coding bottlenecks ease, but judgment, evaluation, experiment compute, energy, chips, labs, security, and permissions remain or intensify.

  7. RDI-07 NOT MET

    Is recursive acceleration independently observed across multiple labs, architectures, and research programs, including negative and failed cycles?

    No common independent multi-lab, multi-architecture dataset contains successful and failed cycles; participating labs did not report dramatic overall acceleration.

  8. RDI-08 PARTIAL

    Are closed-loop gains present in empirical science and hardware as well as software, so the process can acquire new information rather than only recombine existing digital artifacts?

    AI-generated hypotheses have received expert-led wet-lab validation, but autonomous physical experimentation and feedback into successor AI are not shown.

05Observability and measurement bias6 questions
  1. OBS-01 NOT MET

    Are evaluations preregistered, contamination-resistant, budget-matched, scaffold-disclosed, holistic, and reproducible by independent evaluators?

    Independent methods and released logs exist, but preregistration, budget matching, decontamination, scaffold disclosure, holistic scoring, and replication are not standard.

  2. OBS-02 OPEN

    How much do release selection, benchmark retirement, publication bias, survivorship, cherry-picking, and repeated testing inflate the apparent rate of progress?

    Selection, saturation, task retirement, repeated testing, and publication bias inflate perceived progress, while under-elicitation can bias downward; net size is unknown.

  3. OBS-03 OPEN

    Could decisive capabilities or failures be hidden by secrecy, national security, safety policy, competitive incentives, or incident nondisclosure? If so, how should that widen uncertainty without fixing its direction?

    Secrecy can hide capabilities or failures; limited private access suggested modest gaps for participating labs but says little about nonparticipants or classified uses.

  4. OBS-04 METHOD

    Which leading indicators should move before macro outcomes, and what lag structure prevents interpreting missing macro effects too early or excusing them indefinitely?

    Use the sequence capability, workflow use, complementary capital, firm output, sector effects, then macro; fixed 6/12/24/60-month windows prevent premature or indefinite inference.

  5. OBS-05 NOT MET

    Do audit trails link model inputs and interventions to research, firm, sector, and macro outputs strongly enough for causal attribution and error analysis?

    Local logs and experiments exist, but no linked audit chain runs from model/R&D inputs through firms, sectors, and macro output.

  6. OBS-06 NOT MET

    What negative results, near misses, blocked deployments, and human rescues are missing from public datasets, and do conclusions survive plausible missingness bounds?

    Released shadow failures and maintainer reviews reveal missing rescues and negative results; no conclusions are bounded for plausible missingness.

06Economics18 questions
  1. ECO-01 PARTIAL

    Do controlled or credible quasi-experimental studies show at least the pre-specified task-level improvement in quality-adjusted output per hour or cost per successful task?

    Recent field experiments show meaningful efficiency gains, but quality and distributional effects vary by task and by a user's ability to elicit, filter, and verify output.

  2. ECO-02 NOT MET

    Do task gains persist at the firm level after integration costs, review, errors, workflow redesign, demand constraints, and organizational learning are included?

    Recent multi-firm evidence finds task-level speed and quality gains, but no replicated four-quarter firm value-added or TFP result clears the 10% gate after full integration and error costs.

  3. ECO-03 NOT MET

    Is labor productivity or TFP accelerating above pre-AI trends across sectors representing at least 25% of value added for at least eight quarters?

    Recent AI-adoption productivity correlations are not broad-based or securely causal and do not meet the 2-point, eight-quarter, 25%-value-added gate.

  4. ECO-04 NOT MET

    Is aggregate labor productivity and TFP accelerating by at least the macro thresholds relative to a pre-registered counterfactual, with revisions and utilization treated consistently?

    U.S. productivity remains around its long-run rate; no multi-economy eight-quarter TFP acceleration of at least one point is shown.

  5. ECO-05 NOT MET

    Is real output per capita, gross output, or another welfare-relevant output measure accelerating broadly rather than only AI-sector revenues or stock-market valuation?

    Real GDP and output per capita show no broad discontinuity; AI revenue and consumer surplus do not substitute for real aggregate output.

  6. ECO-06 PARTIAL

    How much AI investment is realized, utilized productive capital versus announcements, land or power reservations, duplicated races, replacement, or speculative overbuild?

    AI investment and hyperscaler capex are exceptional, but realized AI utilization, duplication, reservations, depreciation, and output are not isolated.

  7. ECO-07 OPEN

    Do profits, revenue, cost savings, entry, and consumer surplus demonstrate new value creation, or mainly transfers, rents, market power, and capitalization of expectations?

    Revenue and consumer surplus rise while compute cost and capex surge; new value cannot yet be cleanly separated from rents, transfers, subsidies, or expectations.

  8. ECO-08 PARTIAL

    Which tasks are automated, augmented, newly created, or bottlenecked, and do exposure measures predict realized labor demand after accounting for task reorganization?

    Observed use concentrates in digital knowledge work; automation-heavy use correlates with early-career declines, but exposure does not equal automation and task reorganization is endogenous.

  9. ECO-09 PARTIAL

    Are employment, unemployment, participation, vacancies, hiring, separations, hours, and underemployment changing causally and materially in highly exposed groups relative to valid controls?

    No aggregate labor break appears, while early-career high-exposure occupations show concerning declines below a settled causal threshold.

  10. ECO-10 OPEN

    How are real median wages, wage dispersion, labor share, bargaining power, skill premia, and within-occupation pay changing after accounting for composition and inflation?

    Labor share is historically low and distributional exposure is heterogeneous, but no causal AI effect meets the frozen wage or labor-share threshold.

  11. ECO-11 PARTIAL

    Are quality-adjusted prices falling in AI-intensive goods and services, and are cost reductions passed through to consumers rather than captured as margins?

    OpenAI cut Luna's July list price by 80% and Terra's by 20%, but no official final-consumer quality-adjusted basket meets the two-year 10% pass-through test.

  12. ECO-12 MET

    How heterogeneous are adoption and effects across industries, firm sizes, occupations, education groups, demographics, and regulated versus unregulated activities?

    Adoption and effects are strongly heterogeneous by firm size, industry, skill, age, education, income, occupation, and task; breadth for S4 is not met.

  13. ECO-13 MET

    How heterogeneous are capability access, infrastructure, output, labor, and price effects across regions and countries, including spillovers and displacement through trade?

    Access, task mix, infrastructure, exposure, and likely gains differ sharply across countries and regions; U.S. knowledge work cannot represent humanity.

  14. ECO-14 PARTIAL

    What implementation, complementary-capital, regulation, contract, education, and statistical lags are expected, and does a pre-specified J-curve model fit observed diffusion?

    Complementary-capital and intangible J-curves are plausible, but no AI-specific preregistered lag model has passed; frozen horizons prevent indefinite excuses.

  15. ECO-15 NOT MET

    Can observed effects be causally separated from the business cycle, fiscal and monetary policy, post-pandemic normalization, demographics, immigration, energy prices, re-shoring, conventional software, and other innovations?

    Business cycle, policy, pandemic normalization, demographics, migration, energy, trade, and conventional software remain unresolved causal confounders.

  16. ECO-16 NOT MET

    Do conclusions survive half-threshold and double-threshold sensitivity tests, alternate deflators, baselines, windows, exposure measures, and multiple-testing corrections?

    No integrated economic classification survives frozen half/double thresholds, alternate deflators, baselines, windows, exposure measures, breadth, and multiplicity.

  17. ECO-17 OPEN

    Are free digital services, quality change, intangible capital, data, model depreciation, inference as an intermediate input, and consumer time savings measured consistently enough to avoid a productivity mirage or omission?

    Intangibles, cloud boundaries, deflators, depreciation, free services, and nonharmonized adoption materially limit measurement; SNA improvements will arrive slowly.

  18. ECO-18 PARTIAL

    Could concentration, monopoly rents, compute scarcity, intellectual-property control, or unequal access produce extreme private gains without broad social productivity or welfare acceleration?

    AI adoption, investment, compute, and IP are concentrated enough to generate large rents or private gains without broad TFP or welfare acceleration.

07Infrastructure and physical constraints5 questions
  1. INF-01 PARTIAL

    Can useful compute, memory, networking, storage, and inference capacity scale at the rate required by the alleged feedback loop after utilization and redundancy are counted?

    Useful digital capacity and efficiency scale quickly, but agentic compute, utilization, redundancy, memory, networking, and concentrated supply chains remain binding.

  2. INF-02 NOT MET

    Are electricity generation, transmission, interconnection, cooling, water, land, and permitting constraints easing quickly enough, locally as well as globally?

    The July IEA update and June FERC action still show power demand outrunning local grid readiness. Interconnection, transformers, cooling, power, and permitting remain binding constraints.

  3. INF-03 NOT MET

    Can chip design, fabrication, advanced packaging, equipment, materials, and geopolitical supply chains respond without multi-year bottlenecks or single points of failure?

    Advanced logic, memory, packaging, equipment, materials, and geopolitics remain concentrated multi-year dependencies.

  4. INF-04 NOT MET

    Can robots, laboratories, factories, construction, logistics, and maintenance convert digital plans into physical throughput at matching speed and reliability?

    Robotics, laboratories, factories, construction, logistics, and maintenance do not yet convert digital plans at matching speed or reliability.

  5. INF-05 PARTIAL

    Do rebound effects, demand growth, scarce inputs, and induced complexity consume cost/performance gains fast enough to prevent accelerating effective output?

    Energy per task falls while total data-center demand rises, demonstrating rebound; whether rebound prevents useful-output acceleration is unresolved.

08Safety, governance, and control5 questions
  1. GOV-01 NOT MET

    Can developers and operators detect deception, specification gaming, sabotage, data exfiltration, unsafe capability, and correlated agent failure before deployment?

    Safeguard vulnerabilities, cheating, monitor bypass, and assurance gaps prevent reliable predeployment detection of deception, sabotage, or correlated failure.

  2. GOV-02 NOT MET

    Do monitoring, sandboxing, access control, evaluations, incident response, and shutdown remain effective as capability, speed, copies, and tool access scale?

    Defense in depth improves control, but universal finite guardrails are impossible and monitor efficacy at greater speed, copies, and tool access is unproven.

  3. GOV-03 PARTIAL

    Can law, standards, audits, liability, export controls, and international coordination observe and respond faster than the relevant capability and diffusion cycle?

    Institutes, standards, safety frameworks, and coordination expand, but evidence and legal response cycles remain fragmented and often slower than releases.

  4. GOV-04 PARTIAL

    Does concentration in a few labs or states increase control capacity, systemic fragility, race pressure, or the chance that private evidence diverges from public claims?

    Concentration can improve access control and evaluation while increasing information asymmetry, systemic fragility, and race pressure; net effect is unknown.

  5. GOV-05 NOT MET

    Is the trajectory reversible in technical, economic, legal, and geopolitical terms? What observable threshold would mark a loss of meaningful human control?

    No public evidence shows technical loss of meaningful human control; economic and geopolitical lock-in rises but infrastructure and permissions remain human-operated.

09Alternative explanations and falsifiers5 questions
  1. ALT-01 OPEN

    Are apparent exponential trends local fits that will flatten because of diminishing returns, benchmark ceilings, data limits, verification burdens, or increasing problem difficulty?

    Exponential fits are local, saturated, cost-sensitive, and weaker on messy work; recent fast cyber gains update toward acceleration but do not identify supercritical returns.

  2. ALT-02 MET

    Could AI primarily complement scarce experts, raising their leverage while leaving tacit judgment and institutional throughput as the binding bottleneck?

    Current AI primarily complements experts and automates clean components while judgment, agenda setting, review, and institutional throughput remain binding.

  3. ALT-03 OPEN

    Could falling model prices, benchmark gains, and rising capex coexist with weak returns because supply races, commoditization, overcapacity, or Jevons effects dissipate private and social gains?

    Falling costs can coexist with huge capex, free services, competition, duplication, and rebound, producing weak private returns or weak measured productivity.

  4. ALT-04 PARTIAL

    Could safety restrictions, security failures, liability, public resistance, war, trade conflict, or political choice deliberately slow diffusion even if technical capability continues accelerating?

    Security, liability, trade controls, grid permitting, resistance, conflict, and political choice can slow diffusion despite technical gains.

  5. ALT-05 NOT MET

    What observations would distinguish a genuine self-amplifying regime from coordinated marketing, investor narrative, selective disclosure, and anthropomorphic interpretation of agent behavior?

    Independent evaluations prove real progress, but no audited recursive cycles or broad causal outcomes distinguish a self-amplifying regime from stronger marketing/investor narratives.

These are concise lead conclusions from the evidence cutoff. Each independent reviewer also answered all 64 questions, and the complete synthesis records disagreements and data gaps.

METHOD + SOURCES

Definitions came before conclusions.

The questions, stage gates, effect-size thresholds, and monitoring rules were frozen before the new answer research began.

01 How the research was run

The final frozen framework contained 64 unique canonical questions across definitions, capability, autonomy, recursive AI R&D, measurement bias, economics, infrastructure, safety, governance, and alternative explanations. Exploratory questions and analysis from the three predecessor research reports are not included in that count.

An independent lead pass and three separate high-reasoning reviews answered every question. The additional reviews focused on mainstream evidence, contrarian falsification, and economics and measurement. Each pass covered exactly 64 of 64 questions. The reconciliation therefore contains 256 per-question assessments.

All four passes selected S1. Their probability distributions were averaged equally for transparency. Because they share sources, the result is a structured judgment, not a statistical posterior.

02 Core capability and R&D sources
03 Economic and infrastructure sources
04 How the earnings evidence is used

The earnings module uses historical primary filings only to establish each company's launch-era baseline. Every current endpoint comes from the latest filing or earnings release available by July 30, 2026.

Cloud segments, fiscal calendars, capex definitions, useful lives, acquisitions, and backlog measures differ. Values are compared only within companies and are never summed into a synthetic hyperscaler total. Backlog is treated as contracted demand, capex as an input, and revenue as realized sales. None is treated as proof of recursive improvement or economy-wide productivity.

THE BOTTOM LINE

The foothills of a possible singularity.

The best description is rapid AI acceleration with a meaningful possibility of bounded recursive improvement, but very little current evidence for a self-sustaining takeoff or an economy-wide singular regime.

S1: rapid AI-accelerated transition. Meaningful S2 tail. Very low current S3 and S4 probability.

That conclusion is neither complacent nor apocalyptic. Capability is moving fast enough to justify serious work on safety, labor, infrastructure, governance, and measurement. The evidence is simply not yet strong enough to call the feedback loop autonomous, input-controlled, or economy-wide.

Join 5,900+ members From $30/mo