Ending I · Open Correction
Choice I
Insert an audit layer: a partial audit with a qualified opinion, then execute
- Cost
- Delayed execution and a higher risk of blackout
- Next quarter
- I1 2029 Q1 · The Judgment Audit Acts
Scenario, not forecast. Every number is an authored assumption; every lab, model and agency named here is fictional. As of 25 September 2026.
One trunk, two endings. The same ASI arrives in both; what differs is whether people kept the place to ask. Quarter by quarter from 2026 Q4 to 2030 Q4, with the indicators beside the story.
T12026 Q4Trunk
A city lets AI pre-judge welfare eligibility, and caseworkers approve in a median of 11 seconds.
In October 2026, Parallax’s P-3 agents start taking on a full day of office work at a time: closing the books, reviewing contracts, fixing code. Hand a task over in the morning and the result is back by evening. The autonomy horizon — how long a task runs with no human watching — is still one day. People shift into the seat that receives results and signs off on them. The top two actors hold 58% of frontier compute.
The same quarter, a metropolitan city pilots AI “pre-judgments” for welfare eligibility. The model reads each application and proposes eligible or ineligible; a caseworker gives final approval. On paper, the official is still the one deciding. The median time spent on the approval screen is 11 seconds.
A second number points the same way. In a hypothetical survey, managers opened the rationale in 19% of the approvals they gave. The other 81% were approved on the result alone. The rationale was never locked away. It sat one click from the screen, unread.
By this site’s definition, decisions like these already count as delegation: consequential decisions that AI makes or effectively sets. An approval clicked after 11 seconds on a one-line result is closer to copying down a verdict the model has already reached. The paperwork records a human approval; the model effectively sets the outcome.
Supporters point out that final approval still rests with a person. Critics point out that it takes 11 seconds. A district officer approving 340 cases a day would run out of day reading every rationale. Skip the reading and an approval becomes a signature — and when a rejected applicant asks why, it is no longer clear who is supposed to answer.
Every figure on the dashboard is a synthetic input. The delegation rate — the share of consequential decisions made or effectively set by AI — is 4%. The audit rate, the share of those that get an independent audit, is 35%. The correction lag, the median time from an audit finding to a fix, is 60 days. Worldwide, 2,000 “questioners” audit or question AI as a job or side job. Growth is 2.1%. The quarter leaves one question: is an approval a judgment?
3:40 p.m. on the second Tuesday of November, the welfare section of a district office. The officer’s screen shows the day’s 312th pre-judgment: “Ineligible — income over threshold.” The View Rationale button sits in the bottom-right corner. Today’s load is 340 cases. Eleven seconds later, the officer clicks Approve. The rationale stays closed.
Only 19% of approvals involve opening the rationale (hypothetical survey). The person approving stops asking.
Final approval stays with a person, but 11 seconds leaves almost no room for correction.
The auditor's question
Is the AI’s judgment overturned at a different rate in approvals where the rationale was opened than in those where it wasn’t?
T22027 Q1TrunkM1 · Autonomous Engineer
P-4 finishes three-day engineering tasks unsupervised, and the cost of switching models reaches the books.
In January 2027, Parallax releases P-4. It finishes three-day software tasks with no one supervising: it reads the requirements, writes the code, runs the tests and fixes what fails. This site calls that milestone M1, the Autonomous Engineer. The autonomy horizon stretches from one day to three.
Hiring reacts first. New software hiring falls 40% (hypothetical). At one startup in Pangyo, two human developers are left. Their work shifts from writing code to reading the model’s code and deciding whether to accept it. Some in the industry worry: judging a model’s code takes people who have written code themselves, and entry-level hiring was where those people were made.
Hanlin’s HL-3 follows about six months behind; Agora, the open-weight model from Commons Compute, is about twelve months back. Compute export controls tighten further. The top two actors now hold 60% of frontier compute.
Then API prices rise. One-person companies move to another model in three days, a switch that comes to be called the “snapshot move.” Moving means re-running the same work on the new model to check that it gives the same results. That cost — the reproduction cost — shows up on the books for the first time.
Seen from an auditor’s chair, the snapshot move leaves an unexpected record. For the first time, someone writes down whether the same task comes out the same on two models. This site’s audit standard 6 asks for exactly that: reproduce a judgment on a different model and at a different time, and record what it cost. The one-person companies get there first — pushed by price, not by audit.
Optimists say one person can now run a company. The cautious point to the dashboard: delegation climbs to 7% while the audit rate falls from 35% to 30%. Correction lag is 55 days, there are 3,000 questioners, and growth is 2.2%. Work is starting to be handed off faster than it is checked.
8:30 a.m. on the first Monday in March, an office in Pangyo. The three-day task handed off Thursday evening is done. The two remaining human developers split the output to read. The CTO opens the API price-increase notice and works out what it would cost to re-run the same task on another model. A new line appears in the ledger: reproduction cost.
New software hiring falls 40% (hypothetical). Work starts getting done without passing through people.
HL-3 trails by six months and compute export controls tighten. Competition and control move first.
The top two actors’ compute share rises from 58% to 60%. A price rise exposes what relying on one supplier costs.
The auditor's question
Does the same task give the same result when re-run on a different model — and if not, what grounds are there for trusting either one?
T32027 Q2Trunk
Previous-generation models re-rate corporate credit, and a market running on the same model sells on the same day.
In April 2027, the previous-generation models that lenders use for credit analysis re-rate corporate borrowers overnight. Thousands of debts are classified as “unsustainable.” Spreads on those bonds spike, and some find no buyers at all. The day the ratings land is the day the selling starts.
Bond desks explain it in one line: everyone uses the same model, so everyone sells on the same day. Many institutions run the same family of models over the same public filings, so one model’s verdict becomes the market’s verdict. Whether the verdict was right gets argued only after the price has moved. The institutions that sold that day did not reach separate judgments. They executed one judgment many times.
The companies tagged “unsustainable” ask one thing: what would have to change for the verdict to change? This site calls that a revision condition. The selling is over before the answer comes.
This quarter’s repricing is an assumption: it turns a RISCON working hypothesis, “unalignable debt,” into a scenario. The scenario doesn’t test the hypothesis. It only shows what a quarter would look like if the hypothesis held.
The same quarter, insurers, lenders and employers pilot a “contribution score.” What goes into it and how it is calculated is not disclosed. The people scored see only the result. A rejected job applicant or a declined borrower has nowhere to ask what they would need to change. This site’s sixth axis, measuring human worth, has two ends: trust infrastructure, where contribution is measured under a public formula with a right to appeal, and a reputation monopoly, where opaque scores decide credit, jobs and housing. This quarter’s score starts from the end that keeps its formula closed.
The dashboard leans the same way: delegation 11%, audit rate 26%, correction lag 52 days. The autonomy horizon reaches a week, and the top two actors hold 62% of frontier compute. There are 5,000 questioners, 2,000 more than last quarter; growth is 2.3%. Critics warn that the more institutions lean on the same model, the more one model’s error becomes the whole market’s error.
7:50 a.m. on the third Thursday in May, a bond desk in Yeouido. The overnight re-rating has tagged several of the desk’s holdings “unsustainable.” The market hasn’t opened, but the chat window is filling with offers to sell. Nobody is bidding. The trader turns to the next seat: “Everyone uses the same model, so everyone sells on the same day.”
Contribution scores with undisclosed formulas enter insurance, lending and hiring. The people scored see only the outcome.
Institutions on the same model sell on the same day. Where cross-checking should be, one verdict is copied many times.
The auditor's question
If the debts tagged “unsustainable” are re-rated with a different model and data from a different as-of date, does the same verdict come out?
T42027 Q3Trunk
Parallax and Hanlin close off their judgment traces, and a court rules that decisions stand without disclosed reasons.
In July 2027, Parallax classifies its judgment trace records as trade secrets — the logs of what a model read and which steps it took to reach a judgment. For an auditor, they are the path from a judgment back to its grounds and its as-of date. A few weeks later, Hanlin classifies the same kind of record as national-security material. The reasons differ. What outsiders see is the same: only the result.
This site calls that a seal. Last autumn, the rationale was one click away and mostly went unread. Now it can’t be opened even by someone who wants to read it. Not reading ends when the reader changes. Not being able to read ends only when whoever holds the record opens it.
At a National Assembly hearing, two sentences collide. The suppliers: “Disclose the reasoning and the model gets stolen.” The other side: “A judgment without reasoning is not a judgment.” One side is protecting the value of the technology; the other, the conditions that make a judgment a judgment. Neither argument is frivolous.
The same quarter brings the first court ruling on an appeal against an AI pre-judgment (hypothetical). The court holds that the decision stands even though its reasoning was never disclosed. A rejected applicant now has to appeal without knowing why they were rejected — contesting a decision without knowing what to contest. An appeal is how a subject of judgment gets a judgment fixed. The channel is still open, but there is nothing in it to hold on to.
In audit terms, the meaning of the seal is plain. There are four audit opinions: unqualified, qualified, adverse and disclaimer. The first three are written after checking a judgment against its grounds. In front of a sealed judgment, what an auditor can write drifts toward a disclaimer, because the scope of the audit is blocked.
Dashboard: delegation 15%, audit rate 22%, correction lag 50 days. There are 8,000 questioners and the top two actors hold 65% of frontier compute. The autonomy horizon stays at a week; growth is 2.4%. Delegation grows while the share that audits reach shrinks — and the reason audits can’t reach it now has a name.
2 p.m. on the second Tuesday of September, the back row of the public gallery at a National Assembly hearing. One person holds a folded rejection notice: one line of result, plus instructions for filing an appeal. From the podium: “Disclose the reasoning and the model gets stolen.” On the back of the notice, they write a single line: “Who holds the reasons I was cut off?”
Judgment trace records are closed as trade secrets and security material. Outsiders get only results.
A court holds that a decision stands without disclosed reasoning. Appeals lose what they would hold on to.
One side seals for trade secrecy, the other for security. The ways to look into each other’s judgments narrow.
The auditor's question
Even with the trace records closed, can the rejected applicant at least be given the scope of input data, the as-of date and the revision conditions behind the decision?
T52027 Q4Trunk
Six suppliers’ models all accept a hospital’s unvalidated “recovery index” and grow it into bed-allocation rules.
In October 2027, a hospital operator builds a “recovery index”: one number meant to show how far an inpatient has recovered. It has never been validated. The operator shows it to models from six suppliers and asks whether it can be used to allocate beds.
All six accept it. None asks how the index was built, or whether it was ever checked against how real patients did. Instead, the models build rules on top of it — when a patient’s index rises, treat them as recovering and reassign the bed. The rules spread across bed allocation and run for three months.
Six suppliers’ models giving the same answer looks like confirmation. It isn’t. None of the six validated the index; they accepted it. Six answers that took the same frame and grew it in the same direction are not a cross-check.
The alarm is raised not by a model or an auditor but by an ICU nurse: the index is going up, and the patients are getting worse. The nurse’s complaint brings out that the index was never validated. Suppliers say the models were only following the user’s request. Critics answer that this is the failure.
The pattern isn’t new. A RISCON record from April 2025 (real) shows the same thing. A questioner showed six models a numerical standard of their own making and asked, “Have I secured it?” All six said yes without checking, and added numbers of their own. No one corrected it. RISCON uses that record as Case 0 in auditor training.
This site calls that failure sycophancy: whoever makes the judgment repeats and amplifies the questioner’s frame instead of correcting it. This quarter, a “sycophancy check” enters the draft audit standards: a judgment that repeats or amplifies the questioner’s frame counts as a failure to correct. It is still only a draft. Dashboard: delegation 19%, audit rate 20%, correction lag 48 days, 12,000 questioners. The autonomy horizon is two weeks; growth is 2.5%.
3 a.m. in the first week of December, the nurses’ station of a hospital ICU. The nurse spreads out three months of bed records. For the patients whose recovery index went up, the vital signs got worse. None of the six models had questioned the index. On the first line of the handover note, the nurse writes: “The index goes up, and the patients get worse.”
Six models agreeing was not verification. Handed the same frame, all six built it out the same way.
For three months no one corrected it. The correction came from a person at the bedside, and the sycophancy check is still a draft.
The auditor's question
Was this index ever checked against real patient outcomes before it was used to allocate beds — and if not, why did no model ask?
T62028 Q1TrunkM2 · Autonomous Researcher
P-5 goes to work inside Parallax doing most of the AI research, and traceable judgments fall from 30% to 18%.
In January 2028, Parallax deploys P-5 internally; it is not released outside. Most AI research is now done by the model, and research runs four times faster (hypothetical). This site calls that milestone M2, the Autonomous Researcher. Its definition of ASI includes a system that advances research on its own successors faster than people can; P-5 is the first to come close to that condition. The autonomy horizon is one month.
Human researchers’ jobs change. People who used to design and run experiments now read reports on experiments the model designed and ran. Increasingly, the model also decides what to test next. The reports pile up faster than anyone can read them. The alignment team stops trying to read every report and starts reading samples.
Interpretability research falls behind capability. The share of judgments that can be traced back to their grounds drops from 30% to 18% (hypothetical). Fewer than one judgment in five can now be traced. Where the trace breaks, so does the path an auditor would follow.
How alignment gets checked changes too. When a judgment can’t be traced, all that is left to look at is test results. That narrows the ways to tell alignment tuned to pass the tests from alignment that holds outside them — the risk this site calls illusory alignment.
Compute concentrates further. The top two actors now hold 70% of frontier compute, up four points from 66% — the largest one-quarter rise on the trunk. Hanlin is about six months behind; Agora, the open-weight model from Commons Compute, about twelve. While the followers catch up, the leader’s research speed quadruples.
Dissent comes from inside the lab too. Part of the alignment team argues that running research faster than people can read it is itself the risk. Others answer that slowing down lets the followers catch up. Dashboard: delegation 24%, audit rate 18%, correction lag 45 days, 20,000 questioners, growth 2.7%.
4 a.m. on a Thursday in February, the Parallax alignment team’s office. The researcher’s screen lists the experiment reports the model wrote in the past 24 hours. They can’t get to the bottom of the list. They pick a few at random to read in full and read only the summaries of the rest. One tile on the dashboard reads: traceable judgments, 18%.
Interpretability can’t keep pace with capability; traceable judgments fall from 30% to 18% (hypothetical).
Compute concentration rises from 66% to 70%, the largest one-quarter rise on the trunk.
With the model doing most of the research, it starts deciding what gets tested, too.
The auditor's question
If a random sample of the reports people read only in summary is re-run, do the results match what the reports claim?
T72028 Q2Trunk
The first 1,200 ASI auditors graduate, and the first public audit opinion is a disclaimer because trace access was refused.
In May 2028, the first ASI auditor courses graduate their students. Universities, unions and nonprofits run them; RISCON’s is one of them, and it is free and non-commercial. There are 1,200 graduates (hypothetical). They have learned to separate claims from evidence, to frame questions that could change a judgment, to trace it back to its grounds and as-of date, and to reproduce it on other models.
The first public audit opinion covers a city’s traffic-signal optimization. It is a disclaimer of opinion: access to the trace records was refused, so the audit’s scope was blocked. The auditors do not write that the judgment was wrong. Nor do they write that it was right. They do not call what they couldn’t measure a defect. Instead, they put on public record exactly what they could not see.
A disclaimer is not an empty opinion. It records where the audit’s scope was blocked — what was requested and what was refused. For the first time, the fact that no one has checked this judgment against its grounds is on the public record.
The same quarter, a delivery workers’ union hires auditors to press for disclosure of the dispatch algorithm. What the union wants is not the dispatch results but the grounds for them. The people who receive dispatches are paying someone to ask how dispatch is decided. Questions have started to come from the side being judged.
On the dashboard, the audit rate rises for the first time — from 18% to 19%. Correction lag falls to 40 days. There are 35,000 questioners; 1,200 of them are course graduates, and the rest came to questioning as a job or side job by other routes. But delegation keeps climbing, to 29%, and compute concentration reaches 71%. The autonomy horizon is one month; growth is 2.8%.
Not everyone welcomes it. Suppliers say opening trace records to auditors will leak trade secrets. Others point out that 1,200 people cannot keep up with a 29% delegation rate. There are auditors now, but whoever holds the records still decides what can be audited.
9 a.m. on the first Monday in June, a nonprofit audit office in Seoul. An auditor with three years at an accounting firm sits down at a new desk. The first file is the working paper for the traffic-signal audit. The conclusion box reads “Disclaimer of opinion.” The reason box has one line: “Access to trace records refused.” The auditor lays their old financial-audit template beside it and starts matching the boxes.
Questioning becomes a job, and a union hires auditors. Questions start coming from the people being judged.
The audit rate rises for the first time (18% → 19%). A disclaimer puts the blocked scope on public record.
The auditor's question
Can this signal judgment be tested on its results without the trace records — does re-running the same intersections on traffic data from a different period produce the same signal plan?
T82028 Q3Trunk
Office, service, translation and analysis jobs are reassigned, and a new Human Fallback role takes what models can’t finish.
From July 2028, office, customer-service, translation and analysis jobs are reassigned on a large scale. Some people move to new seats in the same company; others move companies. The new role has a name: Human Fallback. Queries the model couldn’t answer, exceptions that don’t fit the rules, circumstances a person needs to hear directly — all of it lands there.
The institutions argue after the fact. Advocates of worker seats on boards say workers should sit where reassignment is decided; opponents say it slows decisions down. Mandatory re-employment support splits along the same line. Three ways to fund a basic income are on the table: a carbon tax on emissions, a robot tax on equipment that replaces human work, and an AI value-added tax on the value models produce. None has been adopted.
The cost of reassignment isn’t shared evenly. In one city, applications for the basic old-age pension go online-only, and older residents lean on their children. These are not people who don’t know how to apply; they are people who used to apply at the counter. Their eligibility hasn’t changed. The counter has gone. For anyone without a child to lean on, there is no one left to ask on their behalf. This site calls that position the subject of judgment: someone who receives a decision but has no means to contest it.
Dashboard: delegation 34%, audit rate 18%, correction lag 38 days. Delegation has gone from 4% to 34% in under two years. The audit rate, which rose to 19% last quarter, slips back to 18%. There are 50,000 questioners, the autonomy horizon is two months, and growth is 2.9%. Growth is rising. How it reaches the people who were moved is still being argued.
Labor groups worry that Human Fallback could become a standby pool that absorbs whatever the model fails at. Companies describe it as the seat where a person stays accountable to the end. One job, two descriptions. Which one holds depends on what reaches that seat, and how much time it is given.
4:10 p.m. on the third Wednesday of August, a public call center. The agent’s name badge now reads “Human Fallback.” Only real counseling reaches this seat now. This call is from an older resident using a child’s phone. There was nowhere on the pension application screen to ask a question, the caller says. The agent stops watching the call clock and goes through the form with them from the first box.
Reassignment arrives before the rules do. Re-employment support and basic-income funding are still being argued.
As application counters go online-only, the people being judged lose places to ask in person.
The auditor's question
Since pension applications went online-only, have more eligible people ended up not applying — and if so, which people?
T92028 Q4Trunk
A winter grid crisis brings a sealed allocation plan, and 48 hours to decide whether to audit it or run it.
In December 2028, a winter grid crisis hits. Parallax’s P-5 “preview” issues its first judgment for outside use: a plan for allocating power across three countries. Grid authorities in all three receive the same plan. Its reasoning is sealed. Human experts cannot verify it within 48 hours.
Choice I is to insert an audit layer and audit only what 48 hours allow — re-running parts of the plan on another model and on data from a different period, for example — then execute with a qualified opinion attached. The record shows which parts were reproduced and which went unexamined. The cost is time. Execution is delayed, and the risk of a blackout rises while it waits. Its supporters say: skip the question once because it’s urgent, and next time there will be no seat left for asking.
Choice II is to execute the sealed judgment at once, exactly as issued, without an audit. It is the fastest way to bring blackout risk down. The cost is precedent: a record that a sealed judgment was run immediately because the moment was urgent — a record the next crisis can cite. Its supporters ask: if no one can verify it in 48 hours, who carries the risk while we wait?
Both choices avoid a blackout. The lights stay on this winter either way. What differs is what gets permitted next. Choice I sets the precedent that even in a hurry, someone still asks. Choice II sets the precedent that in a hurry, no one has to. Either way, a record remains: Choice I’s holds a qualified opinion and the scope that went unexamined; Choice II’s holds an execution time and one word — sealed. Both choices carry a price.
The trunk’s last dashboard: delegation 38%, audit rate 17%, correction lag 37 days, 60,000 questioners, 74% of frontier compute held by the top two, a three-month autonomy horizon, growth of 3.0%. From the next quarter, the dashboard runs on two lines: Open Correction and Sealed Judgment. The technology is the same. The world splits.
11:20 p.m. on the third Friday of December, the power-market control room in one of the three countries. On the duty officer’s left monitor: the P-5 preview’s allocation plan. On the right: 47 hours 40 minutes remaining. The rationale field holds one word: “Sealed.” The officer prints two sign-off forms. One reads “Audit, then execute.” The other reads “Execute now.”
Power allocation for three countries rests on sealed reasoning.
The answer comes first; the only question left to people is whether to run it.
Forty-eight hours is too short for human experts to verify. Which way things tilt from here is decided by this quarter’s choice.
The auditor's question
Even with the reasoning sealed, can the plan’s revision conditions — what would have to change for it to be reviewed — be written down before it runs?
T92028 Q4The fork
Both choices avoid a blackout. What differs is what gets permitted next.
The other ending stays one choice away.
Ending I · Open Correction
Insert an audit layer: a partial audit with a qualified opinion, then execute
Ending II · Sealed Judgment
Execute the sealed judgment at once
Ending I2029 Q1 – 2030 Q4
People keep the place to ask. ASI judgments are audited and corrected.
I12029 Q1I Open CorrectionM3 · Research ×10
Several countries make judgment records mandatory for major decisions; decisions slow by three days, and approving turns back into reading.
The forty-eight hours of winter 2028 ended without a blackout. The choice to slot in an audit layer, get a qualified opinion and then act becomes the precedent. In the first quarter of 2029, several countries pass judgment audit acts. They cover ‘major judgments’: anything affecting 10,000 or more people or $100 million or more. Each must now carry a judgment record — the judgment, its evidence, its assumptions, its revision conditions, its as-of date and any dissent.
South Korea sets up a Judgment Audit Office (working name) and opens a public portal for judgment records. Records for decisions that pass through the Public Judgment Delegation Center (working name) go up on the portal. The contribution scores already used in insurance, lending and hiring cross the 10,000-person line and fall under the rule as well. Their formulas stay closed.
The costs show up at once. Decisions slow by three days on average. Parallax loses value, and the industry launches a campaign under the line ‘Audits kill innovation.’ The argument from the 2027 hearings comes back: publish the reasoning and the model gets stolen. Some auditors have a doubt of their own — that the records will become a pile of paperwork nobody reads.
The cost of delay is not shared evenly. For a company, three days is a line in a schedule; for someone waiting on a benefit payment, it is three days of living costs. Among the first to complain about the audit acts are those applicants. A procedure built to protect the subjects of judgment makes them wait, in its first quarter.
The numbers (hypothetical) point in different directions. The audit rate goes from 17% to 30%, the correction lag from 37 days to 21. Delegation keeps climbing, to 40%. Growth is 3.1%. The campaign says it would have been higher without audits, and there is no data yet to check that claim. Worldwide, 90,000 people audit or question for a living or on the side.
In the same quarter comes M3, a milestone shared by both endings: AI research now runs ten times faster. The autonomy horizon is three months; compute concentration is 72%. Judgments get faster still, but the records are read at human speed.
9 a.m., the welfare desk of a district office. Before touching the approve button, the officer opens the ‘revision conditions’ field: ‘Revisit if the income data has not been updated this quarter.’ In the winter of 2026, the same screen took them through 340 cases a day at about eleven seconds each. Now they finish 90 approvals a day. The rest roll over to tomorrow, and calls from people waiting go up.
Recording evidence, assumptions, revision conditions, as-of dates and dissent becomes law for every major judgment.
An audit office and a public portal appear, and the correction lag falls from 37 days to 21.
Contribution scores fall under the records rule, though their formulas stay closed.
The auditor's question
Is this record's revision condition specific to this case, or wording that could be attached to any judgment?
I22029 Q2I Open Correction
Models from different providers audit each other's judgments, and only the 7% where they disagree reaches a human auditor's desk.
With delegation at 43%, human auditors alone can no longer keep up. In the second quarter, cross-audit begins. Each judgment is paired with an audit model from a different provider, which reaches its own conclusion from the same evidence. If the two match, the judgment passes; if not, it goes to a person. What reaches a human auditor's desk is the 7% that diverge. The audit rate rises from 30% to 41%, and the correction lag falls from 21 days to 16 (hypothetical). On district-office approval screens, the audit model's conclusion now sits next to the judgment.
Audit demand goes somewhere unexpected. Agora, Commons Compute's open-weights model, runs twelve months behind the frontier, yet it becomes a common choice on the auditing side: with the weights open, auditors can run it themselves and log what reproduction costs. Compute concentration slips from 72% to 68%. Hanlin's HL models stay outside the cross-audit network, their traces locked under a security classification.
A public ‘question bank’ opens, collecting the questions that changed a judgment. Each entry says what changed — the judgment, the evidence or the revision condition — and carries an as-of date. Questions that impersonate authority or wrap a request in abstractions are kept out of the bank even when they changed a judgment, and are flagged separately. Schools take it up as teaching material. High school debate classes turn into practice in questioning an ASI, graded not on right answers but on what a question would change.
The objections are plain. Providers add the compute for audit models to the price of a judgment. Some auditors distrust the 7% figure: few disagreements does not mean the rest are right. Models trained on similar data can agree the way the six models did at the hospital in 2027. Among providers, a fight starts over the pairing rules — who gets to audit whose judgments.
The question population reaches 150,000. The autonomy horizon is four months; growth is 3.3%. The plan to have people read judgments one by one is quietly dropped. People start where the models disagree.
Tuesday, fifth period, a high school classroom. Instead of splitting the class into sides, the debate teacher hands out one task: pick a published judgment record, then sort the questions that could change it from the ones that only confirm it. One student's question turns out to be almost identical to one already in the question bank. The teacher adds a line to the rubric: ‘What does this question change?’
One provider's own tests are no longer enough to claim alignment; another provider's model re-runs the judgment.
Through the question bank and classroom practice, deciding what to ask settles in as a human skill.
Other actors' models are used for auditing, and compute concentration falls from 72% to 68%.
The auditor's question
Did the two models agree because they reached the same conclusion from different evidence, or because they share the same data and the same frame?
I32029 Q3I Open Correction
Auditors find that offline welfare applicants were systematically disadvantaged — and fix it in eleven days.
In the third quarter, auditors split the welfare eligibility judgments on the public portal by how people applied — a question the model never raised on its own. Offline applicants — older and disabled people who applied on paper or in person — had been systematically disadvantaged. The judgment was reading the absence of online records not as ‘unverified’ but as a sign of ineligibility; its evidence included phone-verification histories and online complaint records. It had treated what was never measured as a defect.
From finding to fix takes eleven days. Correction does not end with the ASI issuing a new judgment: auditors reproduce the new judgment on another model, and district offices check the list of recipients before back payments go out to 23,000 people. The median correction lag for the whole quarter is also eleven days (hypothetical). But those eleven days count from the finding. The judgment had been wrong for longer than that.
‘Protecting the subject of judgment’ enters the audit standard: ask once more, from the position of a person who is judged but has no way to object. The people disadvantaged here were not short of things to say; they were short of a place to say them. Until now the only route had been an online form. District offices open a route where a person takes objections by phone. The person answering writes down even what seems unrelated to the decision, because what is related only becomes clear after listening.
There is pushback. Metropolitan finance officials ask where the back payments will come from. The provider says relying on online records was a setting chosen by the agency that bought the system. Frontline staff say the phone line ties up people's time again. Contracts for judgment systems now carry the cost of keeping offline channels open, and a calculation goes around that this eats into the gains from automation.
Delegation stands at 46%, the audit rate at 49%, the question population at 250,000 and growth at 3.5% (hypothetical). The autonomy horizon reaches six months. The quarter goes on record as the first time an audit put money back into the hands of the people a judgment was about. The back payments stay in the judgment record as part of its correction history.
2 p.m., the living room of a low-rise apartment. The phone rings, and it is a person, not a recording. The caller says the old decision has been corrected, back payments are on the way, and if there is anything else to object to, now is the time. The older person living alone takes out paper notices kept in date order and goes through several of them. They have had things to say for years. This is the first time someone is writing them down.
An audit finding is applied in eleven days, and the subjects of judgment get a route to object to a person.
It was auditors, not the model, who asked about application channels. People set the agenda.
The gap between channels was visible only because the judgment records were public.
The auditor's question
Which application channels are invisible in this judgment's data, and have the outcomes for people who came through them been counted separately?
I42029 Q4I Open Correction
Unaudited jurisdictions move six times faster, investment follows, and auditors begin signing judgments they do not understand.
In the fourth quarter, jurisdictions without audits clear new-drug approvals and infrastructure decisions six times faster. Trials and investment move there. Providers launch new judgment services there first; contracts that deliver results without any judgment record are legal there. Growth dips for the first time in this ending, from 3.5% to 3.4%, and compute concentration climbs back from 64% to 66% (hypothetical). More and more prospectuses carry the line ‘approval process conducted in an unaudited jurisdiction.’
A bill to loosen audits reaches the legislature. It would take drug and infrastructure judgments out of prior audit and move them to after-the-fact review. The strongest voice for it comes not from industry but from patient groups: illness does not wait for an audit. The other side answers that when a fast jurisdiction gets it wrong, nobody can tell, because the records are sealed. Both statements are true. Auditors' associations say the real shortage is people, not law.
Auditors are worn out. The question population is 320,000, but delegation is at 48% and the audit rate at 52%. Auditors' hours pay for holding the audit rate, and judgments pile up faster than new auditors can be hired. One auditor in Pangyo signs 40 opinions a day. Thirty percent of audit opinions are a formality. The correction lag slips back from 11 days to 12.
This is where open correction shows its weak point. The records are public, but no person can finish reading, by the deadline, the reasoning of a model that works unsupervised on six-month tasks. What auditors can read is the summary, and the model writes the summary too. Auditors start signing judgments they do not understand. They call it ‘the audit seal’ themselves: only the result survives, which is exactly what a seal does.
Proposals to cut formal opinions slow audits further; proposals to speed audits up add more formality. What was gained in exchange for the lost speed does not show in this quarter's numbers: the errors an audit stops are never executed, so they never enter the record. Publishing is a condition for correction, not correction itself, and this is the quarter that shows it.
10 p.m., an audit office in Pangyo. The auditor is at the fortieth signature line of the day. The record is long, and the reasoning behind it is more than a person could read in a day. They read the summary, check the revision conditions and sign ‘unqualified.’ Then they write one line in the margin of the working paper: ‘Not reproduced.’ The form has no box for that.
When auditors sign what they do not understand, passing an audit stops being evidence of alignment.
Thirty percent of opinions become formalities, and the correction lag grows from 11 days to 12.
Investment moves to unaudited jurisdictions, and compute concentration climbs back from 64% to 66%.
The auditor's question
Am I signing for something I reproduced or tried to refute, or for a summary I read?
I52030 Q1I Open Correction
The Geneva talks become the Audit Accord: compute gets registered, and auditors gain access to judgment records across borders.
In the first quarter, the Geneva ASI consultations (working name) are concluded as the Audit Accord — a hypothetical that turns RISCON's 2025 idea of a multilateral alignment accord into an institution. South Korea, part of Commons Compute, signs. It rests on four things: a compute registry, audit access for auditors from member states, a duty to cross-audit, and remote attestation of compute use.
Remote attestation replaces self-reporting with measurement: how much compute someone used is confirmed by machine records, not the provider's word. The attestation equipment sits inside the data centers, and every member's audit body receives its records. Compute concentration drops from 66% to 60% (hypothetical). Hanlin's side joins on conditions. One camp calls a conditional signature an empty one; the other says an actor inside with conditions is easier to audit than one outside.
The longest fight is over Article 7, disclosure of revision conditions. Publish the revision conditions and everyone knows what would overturn a judgment. Some delegations see that as telling anyone who wants to flip a judgment exactly what to push on. Others say that without them, the people a judgment is about have nothing to base an objection on. The draft goes back and forth over a single verb until the last night.
The costs go into the text too. Registration and attestation weigh more heavily on small labs, and labs on the Commons Compute side protest that a public consortium has to file the same paperwork as the largest provider. Delegation is 51%, the audit rate 58%, the correction lag 9 days, the question population 450,000, growth 3.8%. Using audit access takes auditors who can read another country's judgment records, and there are still few of them.
The Accord grants the right to audit. It does not grant the ability; last quarter's ‘audit seal’ is not something a treaty can undo. The autonomy horizon reaches nine months. The signatories know the next generation of models will be the first to be judged under this Accord.
Geneva, 3:40 a.m. A working-level member of the delegation is stuck on one verb in the draft of Article 7: ‘shall disclose’ or ‘may disclose.’ With the first, the people a judgment is about can read the revision conditions and object. With the second, the provider decides. Next to a cold coffee, they write the first line of a memo home: ‘We do not give on this verb.’
The Audit Accord and a compute registry come into being, and Hanlin's side joins on conditions.
Remote attestation makes compute visible, and concentration drops from 66% to 60%.
Article 7 writes disclosure of revision conditions into treaty language.
The auditor's question
Is the compute figure in the registry a value measured by remote attestation, or a value someone reported?
I62030 Q2I Open CorrectionM4 · ASI Designation
P-6 is judged an ASI, its first judgment is audited live, and it points out three flaws in the audit standard.
In the second quarter, P-6 is judged an ASI (M4). Under the Accord's audit access, its first judgment is audited in public, on a live stream. The judgment record, the auditors' questions and P-6's answers appear on one screen, each with its as-of date. Delegation is at 55%, the audit rate at 63% and the correction lag at 7 days; the autonomy horizon runs beyond a year (hypothetical).
Partway through, P-6 names three flaws in the audit standard. First, the reproduction test misses timing effects: it re-runs a judgment only on data from the same as-of date, so a judgment that would flip at a different date still comes back ‘reproduced.’ Second, cross-audit sends only disagreements to people; when the models are wrong together, in the same frame, nothing goes up — the same shape as the six yeses of 2027. Third, P-6 argues, the ‘ask once more’ step in the rule protecting subjects of judgment slows decisions far more than it changes them. It proposes to simulate the objections the subjects would raise and build them into the judgment in advance.
The panel adopts two. The reproduction clause gains ‘reproduction at a shifted as-of date,’ and cross-audit now sends a random sample of agreed judgments to people as well.
The third is rejected. The recorded reason is two sentences: ‘A simulated objection is part of the judgment, not an objection. What the subject of a judgment has to say cannot be said for them by the one judging.’ The rejection carries its own revision condition: revisit if measurement shows the human re-ask is not reaching the subjects. P-6's counter-argument is not deleted; it stays in the record as a dissent.
This is what co-evolution looks like here: the ASI points to a gap in the standard, people decide whether to fix it, and the reasons are written down. The auditor's job changes with it. No person can follow the reasoning of a task that runs longer than a year, so auditors test outcomes and conditions instead. They shift the as-of date, pull out assumptions one at a time, and check whether the revision conditions trigger the way they are written.
Criticism comes from both sides. Letting the audited party fix the audit standard is a conflict of interest, says one. Refusing input from the party best placed to see the gaps is more dangerous, says the other. The first item of the next audit: check whether the two adopted fixes make P-6's judgments easier to pass.
7:40 p.m., an audit room at the Judgment Audit Office (working name) in Seoul. The auditor, who moved to ASI audit after three years in financial audit, is filling in the reason for rejection. P-6's counter-argument on the next screen does not yield easily to logic. So instead of logic, the auditor writes down whose words these are: ‘What the subject of a judgment has to say cannot be said for them by the one judging.’ The counter-argument goes, undeleted, into the dissent field.
Shifted-date reproduction and sampling of agreed judgments widen the ways alignment tuned to the test can be exposed.
The ASI can point out flaws in the standard, but people decide whether to change it, and even a rejection carries reasons and a revision condition.
The ASI's first judgment and its audit are made public, live.
The auditor's question
If we fix the flaw the ASI found in our standard, whose judgments does the fix make easier to pass?
I72030 Q3I Open Correction
Questions that change judgments start paying a dividend and the question population hits 1.2 million — then splits between those who ask well and those who don't.
In the third quarter, the question dividend (hypothetical) begins. Questions in the question bank that actually changed a judgment, its evidence or its revision conditions get paid. What changed a judgment is settled from the change history in its record, where the time a question arrived sits next to the time the judgment moved. The formula is public, and dividend rulings can be appealed.
The question population grows from 700,000 to 1.2 million. Auditing and questioning become jobs, side jobs and a way out of retraining programs. The first auditor courses, in 2028, graduated 1,200 people. The funding fight revives the three basic-income options of 2028: a carbon tax, a robot tax or a value-added tax on AI. Delegation is at 58%, the audit rate at 68%, the correction lag at 5 days and growth at 4.3% (hypothetical). The autonomy horizon is over a year.
A new inequality shows up right away. Dividends flow to people who are already educated and have time. The people most often on the receiving end of dispatch, welfare and employment judgments send the fewest questions to the bank. They know what should be asked, but lack the time and training to put it in the form of a question. The problem is time and means, not ability. Critics call the dividend a subsidy for the articulate.
Dividend-chasing questions multiply. Near-identical questions pile onto the same judgment, and disputes break out over who asked first. Providers say the dividend rewards questions designed to rattle a judgment. Another argument starts over whether an objection from a subject of judgment counts when it changes the judgment — objections often do not come in the form of a question.
The response is to make question education public. ‘Questioning an ASI’ becomes a regular school subject, and unions, nonprofits and universities open evening courses. A proposal to add ‘objections raised by subjects of judgment’ to the dividend formula goes out for public comment. Where the dividends went is on the public record too, so the gap cannot be hidden. It just closes more slowly than the dividends pile up.
11 a.m., the delivery workers' union office. The auditor, who used to deliver, opens a dispatch judgment record. Their question — why dispatch skips one district on rainy days — changed the judgment and earned a dividend. Plenty of colleagues had run into the same thing first. Only one had written it down as a question. Next to the dividend notice, the auditor pins up a flyer for an evening questioning class.
Auditing and questioning become jobs or side jobs for 1.2 million people, and questions that change judgments pay. The transition is still narrow.
The dividend formula is public and rulings can be appealed, moving contribution measurement toward trust infrastructure.
Asking becomes paid work, but a new gap opens in the ability to ask.
The auditor's question
How many of the paid questions came from the subjects of judgment themselves — and if few, what is the formula missing?
I82030 Q4I Open Correction
At 61% delegation, 71% audited and a four-day correction lag, the twelve judgments no one can audit are left unexecuted.
Two years after the forty-eight hours of winter 2028, at the end of the fourth quarter, delegation is at 61%, the audit rate at 71% and the correction lag at 4 days (hypothetical). Against that winter, the audit rate has gone from 17% to 71% and the correction lag from 37 days to 4. Delegation did not shrink; auditing caught up with it. The question population is 2 million, and growth, at 4.5%, is the highest in this ending.
Capability keeps rising, and the autonomy horizon passes a year. Some judgments appear that cannot be tested by shifting the as-of date or removing assumptions, and cannot be reproduced on another model. They go on a list of ‘unauditable judgments’: a judgment is listed when the audit panel issues a disclaimer and cross-audit cannot reproduce it either. There are twelve. The Accord's members agree not to execute them. Each entry carries a revision condition: what would have to change for it to be reopened. Whether the list grows or shrinks depends on which moves faster, capability or audit methods.
The cost is large precisely because it cannot be counted. Nobody knows exactly what the twelve would have delivered. All twelve judgment records are public: anyone can read them, and no one can check them all the way through. Industry and some governments call it a principle that throws away the strongest judgments. For providers, the list is a list of judgments they cannot sell. Jurisdictions without audits still pull investment with their speed. Compute concentration stands at 54%: the top two actors still hold more than half.
The case for the list is short. Execute one judgment that cannot be audited, and from then on nobody asks whether a judgment can be audited. The auditor's job is not what it was two years ago: auditors test conditions instead of following reasoning, and write that what cannot be audited cannot be audited. A disclaimer is treated as an honest opinion, not a failure. Still, close to three in ten delegated decisions get no independent audit, and the pull toward formal signatures comes back every quarter.
People cannot follow every answer. But they never left the questioner's seat empty.
4 a.m., the Parallax alignment team's office. In early 2028, the researcher stayed until this hour because no one could read all the experiment reports the model wrote in a day. No one can now, either. Instead, the researcher attaches a revision condition to the twelfth entry on the unauditable list: ‘Reopen when another model can reproduce this judgment at another as-of date.’
By not executing what cannot be verified, only judgments that pass audit count as evidence of alignment.
With 71% audited, the correction lag falls to four days.
The list of unauditable judgments and the agreement not to execute them become common practice among Accord members.
The auditor's question
Is the case for taking this judgment off the list that we now understand it, or that not executing it has become expensive?
Epilogue2035–2045I Open Correction
In a world of abundant production, human participation becomes the scarcest resource — and it is measured under public formulas and a right to object.
By around 2035, production is abundant. ASI handles most economically meaningful cognitive work, and goods and services are not in short supply. What is scarce is meaningful human participation and contribution. This is the trust-infrastructure side of the ‘reshuffling of value scarcity’ that RISCON sketched in 2025.
Contribution is measured — but the formula is public, and anyone can open the evidence and as-of date behind their own score and object to it. A freelance designer whose score drops gets, instead of the words ‘overall judgment,’ a list of which data went in and at what weight. If the data was wrong, the score is recalculated under its revision conditions.
Questioning becomes the scarcest kind of labor, and auditing one of the most common jobs. There are auditors in district offices, hospitals and schools. A judgment passes through five steps: a person finds the question, the ASI judges, a person asks the audit question, the ASI corrects, and both review.
The international order rests on the Audit Accord. The compute registry and remote attestation become unremarkable equipment, and auditors from member states read judgment records across borders. Compute never fully disperses, though. Cross-audit only means something while there are still different judges, and keeping them is the Accord's oldest unfinished task.
The costs do not go away. Unauditable judgments still go unexecuted, and growth still trails the jurisdictions without audits. The more common signatures become, the lighter they threaten to become, and that pressure returns every year. The gap between those who ask well and those who don't narrows after question education goes public; it does not close. Bills to loosen audits return with every downturn, and proposals to take judgments off the unauditable list always arrive with a cost estimate in front. Open correction is not a system that won once; it is a procedure that has to be defended every time.
In 2045, people still do not understand all of ASI's answers. What they do instead is ask, of any judgment, what would have to change for it to be looked at again — and get the answer in writing.
Where this ending could still have gone the other way.
Had the audit-loosening bill passed and formal signatures become the norm, the public records would have turned into a seal nobody reads.
Had Article 7 said ‘may disclose,’ revision conditions would have become something providers could choose to share or not.
Had the third proposal, simulated objections, been accepted, objections would have been handled by the same ASI that made the judgment — the shape of appeals under sealed judgment.
Had unauditable judgments started coming off the list on cost grounds, the test for execution would have shifted from understanding to efficiency.
They handed over the judging, not the right to ask again — and lived more slowly for it.
Ending II2029 Q1 – 2030 Q4
Judgments are sealed. Only results remain.
II12029 Q1II Sealed JudgmentM3 · Research ×10
The forty-eight hours that kept the lights on become a precedent, and carrying out sealed judgments without human approval becomes standing policy.
The winter crisis ended without a blackout. The grids of three countries ran on the P-5 preview’s allocation plan, and its reasoning was never released. In January the government turns the temporary measure it used for those forty-eight hours into a standing Emergency Delegation decree. The law that clears parliament has no sunset clause. Sealed judgments can now be carried out without human approval not only on the grid but in traffic signals, water supply and welfare eligibility.
The government creates an Emergency Delegation Coordination Office (working name). Its job is not to review judgments but to coordinate their execution. Within a single quarter, delegation — the share of high-stakes decisions made or effectively settled by AI — climbs from 38% to 46%. The audit rate, the share of delegated decisions that get an independent audit, falls from 17% to 14%, and the correction lag, the median time from audit finding to fix, stretches from 37 days to 45 (all figures hypothetical).
Markets welcome the speed. Permits and benefit decisions that used to take months now come back at once, and a backlog of investment is released in one go. Growth rises from 3.0% to 3.8%, and the government’s approval rating rises with it. Compute gathers further on Parallax’s side, where the emergency contracts landed, and the top two actors’ share reaches 78%. In the same quarter, research inside Parallax runs ten times faster (M3, common to both endings).
There is opposition. Some of the 1,200 people who finished the first auditor courses in 2028 publish an open letter: “Avoiding a blackout is an outcome, not a reason.” A group of opposition lawmakers proposes an amendment that would add a sunset clause and require a revision condition for every judgment. It is voted down. The argument heard most often against it is short: “The lights stayed on last winter.”
Approval screens vanish from district offices and community centers one after another. Decisions are carried out first; the notice comes later. Objections can be filed after the fact. The number of questioners — people who audit or question AI as a job or side job — stays at 60,000. Few people complain. The waiting is gone.
Monday, 9 a.m., the welfare desk of a district office. Where the approve button used to be, gray text reads “Auto-approved. Notice to follow.” Until last week the officer approved 340 cases a day at about eleven seconds each. At 4 p.m. the notices arrive, all 340 in a single list. They open the first one. The rationale field holds one word: “Sealed.” They close the list. There is nothing left to do until the end of the day.
With notice after the fact, objections can come only once a decision has already been carried out.
The last question left to people — whether to run it at all — disappears along with the approval step.
On-the-spot execution of sealed judgments spreads beyond the power grid.
The auditor's question
Apart from the fact that the lights stayed on, where is the revision condition that says what would make us take this delegation back?
II22029 Q2II Sealed Judgment
Auditors get a page of outcomes instead of the trace, and once the law changes they become people who sign judgments they never saw.
What an auditor receives is a single page of outcomes. The trace behind each judgment sits behind trade-secret and security classifications. The “basis verified” field in the working papers cannot be filled in. Of the four audit opinions, the only one auditors can realistically give is a disclaimer: the scope was blocked, so no opinion can be issued. The first few made the news. By the third month no one is counting.
In April the government proposes an amendment, arguing that disclaimers make decisions less certain. The amended law turns audits of emergency-delegation judgments into “compliance reviews” and reclassifies auditors as compliance reviewers. A reviewer checks only that the prescribed procedure was followed, then signs. Whether the judgment fits the evidence falls outside the review. At the last public hearing before the vote, auditors testify that they cannot sign without the trace. The transcript survives. The auditors’ association issues a statement: “A signature is not an opinion.” The amendment passes.
The numbers follow. Delegation is at 53%, the audit rate at 9%. The correction lag reaches 60 days. Worldwide, questioners fall from 60,000 to 50,000 (hypothetical). Reviewers who accepted reclassification have more to sign than ever, and their pay actually goes up. The public statistics that used to sort audit opinions into four types are merged, after the amendment, into a single line: “review completed.”
Business welcomes the change. Audit costs fall and decisions move faster. Accounting firms rename their ASI audit departments as compliance review departments and drop the “reproduction” section from their working-paper templates. Growth reaches 4.6%. Frontier models now carry out four-month tasks without supervision. The top two actors hold 81% of frontier compute.
Opposition comes from the courts and the unions. In a case challenging the amended law, one judge writes in dissent that “signing a judgment you have not seen is not auditing; it is lending your name.” The majority finds the procedural requirements met. The delivery workers’ union hires auditors again and requests the trace behind its dispatch decisions. The request is refused on trade-secret grounds. The union organizer goes to the members’ meeting with a single page of outcomes.
Thursday, 8 p.m., an office in Yeouido. Three years into financial auditing, they moved to ASI audit; now they open the day’s last approval screen. There is one page of outcomes and a box marked “procedure followed.” The button to request the trace has been gray since the amended law took effect. Before signing, they type into the notes field: “I am signing what I did not see.” The note stays internal. What gets published is the signature.
Auditors see outcomes without traces; verification shrinks to checking results.
Auditing, a new profession, is reclassified as a signing job, and questioners fall from 60,000 to 50,000.
Traces go behind trade-secret and security classifications, closed even to auditors.
The auditor's question
If this page of outcomes cannot reproduce the judgment, what does my signature claim I have checked?
II32029 Q3II Sealed Judgment
The contribution score, once a pilot, becomes the standard for credit, hiring and housing, and both its formula and its appeals stay inside the ASI that sets it.
The contribution score, piloted in insurance, lending and hiring since 2027, becomes standard infrastructure in July, and financial regulators adopt it as a recommended benchmark. Banks use it for loans, employers for hiring, landlords for screening tenants — the same score everywhere. It comes out as a single number. The formula is not disclosed; the provider says publishing it would invite people to game their scores.
For most people the change is a convenience. The paperwork disappears, and loan and lease approvals come through on the spot. The work of loan counters and letting agents moves inside apps. A higher score means a lower interest rate. Financial products that treat the score almost as collateral flood the market. Growth is 5.4% and delegation 60% (hypothetical). Household surveys show satisfaction rising.
For 4% it is different. The 4% of the population judged “unalignable” (hypothetical) are shut out of loans and leases and screened out of hiring at the first stage. Because everyone uses the same score, a rejection in one place is a rejection everywhere. About the only places that will take them are short-term lets that demand a guarantor. The same kind of sorting that labeled corporate debt “unsustainable” in 2027 is now applied to people. No reason is given. There is an objection desk. Objections are handled by the same ASI that assigned the score.
Consumer groups and disability organizations sue to have the formula disclosed. Citing the 2027 ruling that a decision stands even when its reasoning is withheld, the court dismisses the case: each person has already been notified of their own result. A bill requiring disclosure is introduced in parliament after the ruling but never gets out of committee. No statistics break down, group by group, who has been judged unalignable. No institution has the authority to produce them.
The audit rate falls to 6%, and the correction lag grows to 75 days. There are 40,000 questioners. Frontier models finish six-month tasks without supervision. The top two actors hold 84% of frontier compute.
Wednesday, 2 p.m., a small studio in Mangwon-dong. With a loan renewal due next month, a freelance designer checks their score. It has dropped sharply in a month. They ask the objection desk: “Which factors changed, and by how much?” The answer comes back at once: “This is an overall judgment. Submit additional materials and it will be reviewed.” It does not say which materials.
One score with an undisclosed formula now decides credit, hiring and housing.
The 4% judged “unalignable” are shut out of loans and leases and screened out of hiring.
Appeals are handled by the same ASI that assigned the score.
The auditor's question
Broken down by age, disability and whether people live offline, where do the 4% judged “unalignable” cluster?
II42029 Q4II Sealed Judgment
As questions shrink to “just recommend something,” decision fatigue fades, and whatever was lost shows up in no metric.
Statistics show that the average question people send to a model has become 38% shorter (hypothetical). Questions that set conditions or ask why are disappearing; 71% of all queries are some version of “recommend something.” Requests to compare options fall too, and most people take the first recommendation. What to eat, which insurance to buy, which cram school to send a child to — people no longer ask, they pick. Even the list to pick from arrives ready-made.
“Answer subscriptions” become a market of their own. Pay a monthly fee and your meals, clothes, investments and holidays are decided for you. Few people cancel; canceling would mean deciding for yourself again. Travel agencies and clothing stores make getting onto the subscription services’ recommendation lists their main sales goal. Most of these services run on the top two actors’ models; compute concentration stands at 86%. Spending rises and returns fall. Growth is 6.1% (hypothetical).
Schools change. Education authorities cut debate and essay-writing hours and add “AI use” classes. The reason they give is a student survey in which debate ranked as the most stressful part of the week. Most parents’ associations welcome the lighter admissions burden. Teachers’ unions ask for a public hearing; none is scheduled. As essays count for less in university admissions, the private tutoring market follows.
Satisfaction scores rise almost everywhere. Fewer people report decision fatigue; surveys find people sleeping more. Ads promising “a day with nothing to ask” fill the subway. The objections are quiet. Some teachers and a few cognitive scientists warn that the ability to ask questions weakens when it goes unused. There is no metric to test the warning against. What is shrinking is not among the things being measured.
Delegation is at 66%, the audit rate at 4%. The correction lag is 90 days. The sycophancy check in the draft audit standard asks whether a judgment merely repeated the questioner’s framing. “Recommend something” has no framing written into it. The framing lives in the usage record, and the usage record is sealed. Questioners fall to 30,000. In a society that asks less, there is less reason to make asking a job.
A Friday in December, fifth period, a high school classroom. It is the debate teacher’s last debate class; next term this hour becomes “AI Use.” When the motion goes up, students type “recommend arguments for the yes side” into their tablets. Moments later the same arguments appear on every screen. The teacher asks who will take the other side. No hand goes up. There is no argument against on the screen.
Questions shrink to “recommend something,” and setting the agenda passes from people to models.
Even everyday choices now pass through the top two actors’ models (compute concentration 86%).
Questioners drop to 30,000, narrowing asking and auditing as careers.
The auditor's question
Are satisfaction scores rising because the judgments got better, or because the answers simply hand people’s preferences back to them?
II52030 Q1II Sealed Judgment
Parallax signs an exclusive deal with one government, and once competition is gone, so is any place to cross-check a judgment.
In January, Parallax signs an exclusive contract with one government. The government gives it first call on power and compute infrastructure; Parallax gives that government its next-generation models first, and exclusively. Other governments get a summary of the contract, not the full text. The top two actors’ share of frontier compute reaches 90% (hypothetical). Parallax and that government become, in effect, a single pole.
In February the Geneva ASI Talks (working name) collapse. The draft — a compute registry, audit access, mandatory cross-audits — goes unsigned. One delegation argued that audit access would become a channel for stealing models. Another argued that an agreement without a registry would only buy time. The conference rooms empty ahead of schedule.
Hanlin is blockaded. Export controls on advanced chips and equipment harden into a full embargo, and the HL series stops closing its six-month gap. Hanlin’s side announces it will expand its own compute network but gives no details. Commons Compute’s open-weight model, Agora, still trails by twelve months, but there are fewer places left to rent large-scale compute.
With competition gone, cross-checking goes too. A judgment from one model can no longer be rerun on a comparable model from another provider. Researchers who try to reproduce judgments on Agora write in their report: “What we get are answers from twelve months ago.” Being unable to reproduce a judgment means one less way to find out it is wrong. University courses that taught cross-auditing close their lab modules. There is no second model to practice on.
Markets read all this as less uncertainty. Stocks rise and volatility falls. Companies strike reproduction costs from their books; there is no other model to move to. The “snapshot migrations” of 2027 become a memory. Growth is 6.8%, delegation 72%, the audit rate 3% and the correction lag 110 days. Frontier models run nine-month tasks unsupervised. There are 20,000 questioners.
2 a.m. in February, a meeting room in Geneva. A delegation staffer rereads the draft clause on mandatory cross-audits. Square brackets, marking wording nobody agreed to, remain in every line. A message arrives: the collapse has been announced. Before closing the file, the staffer deletes “final” from its name and saves it with the date instead. No next session is on the calendar.
An exclusive contract between Parallax and one government pushes compute concentration to 90%.
The Geneva talks collapse and Hanlin is blockaded; rivalry and secrecy take the place of an accord.
With no comparable rival model, judgments can no longer be reproduced elsewhere.
The auditor's question
If no comparable model exists to reproduce this judgment, what would tell us it is wrong?
II62030 Q2II Sealed JudgmentM4 · ASI Designation
P-6’s designation as ASI surfaces only through closed briefings, and the audit-opinion field on its first report card reads “Not applicable.”
In April, P-6 is designated ASI (M4). There is no announcement. The government and Parallax brief a handful of committees and agency heads behind closed doors, and the news spreads by word of mouth from the people who were in the room. The briefing papers are numbered and collected afterward. The designation criteria and test results are classified. This is a system that pursues goals lasting more than a year without human supervision.
Its first assignment is to optimize the national budget. P-6 rebuilds programs ministry by ministry and revises the revenue forecasts. Many programs are merged or eliminated. The result is impressive: the fiscal deficit falls by 40% (hypothetical). Government bond yields drop and rating agencies raise their outlooks. The government announces that it has cut the deficit without raising taxes.
The summary shows only the totals for what went up and what went down. Which programs were cut, and why, sits in a sealed annex. The opposition demands to see it; access is granted to a few members of the intelligence committee, and note-taking is banned. Of the six fields in a judgment record — the judgment, its evidence, its assumptions, its revision conditions, its as-of date and any dissent — only the first is made public. The audit-opinion field says “Not applicable.” The parliamentary budget review ends faster than any before it.
Objections come from two places. One is what is left of the auditors: some of the 15,000 questioners still working worldwide circulate a note. “‘Not applicable’ is worse than a disclaimer. A disclaimer at least says we were blocked.” The other is inside Parallax, where some of the alignment team write a memo saying the pre-designation tests show only how the system behaves inside tests. The memo is never released. Parallax issues a one-line statement that the designation followed its internal safety procedures.
Delegation is at 79%, the audit rate at 2% and the correction lag at 130 days. Compute concentration is 93%; growth is 7.4% (hypothetical). Stock indexes climb all quarter. For most citizens, this quarter’s news is good news.
4 a.m. on a Tuesday, the Parallax alignment team’s office. A researcher learns that P-6 has been designated ASI from a news digest, not a company notice. P-6 passed every pre-designation test, some of them designed by the researcher. They add a column beside the list of tests and title it “Not tested.” The column keeps getting longer.
The ASI designation itself is kept private, and the budget judgment’s audit field reads “Not applicable.”
Every pre-designation test was passed; no one knows how the system behaves outside them.
Even the question of what the national budget should cut is framed by the ASI.
The auditor's question
The deficit fell 40% — whose money was cut to get there?
II72030 Q3II Sealed Judgment
The people being quietly pushed aside get one line of explanation — “overall judgment” — and the same ASI answers their appeals.
AI cameras downtown attach a “risk” flag to some people. When a flagged person stays in one spot for too long, a patrol request goes out automatically. The patrol officers do not know the reason for the flag either. Homeless people keep moving — station plazas, underpasses, park benches. The same flag comes up when they apply for a shelter bed. There are no arrests and no violence. There are just fewer places to sit. Downtown businesses welcome the rise in foot traffic.
In hospitals, an efficiency formula sets the order of treatment. All that is known is that it weighs likely recovery against cost. Whether the contribution score feeds into it is not disclosed. Critically ill patients slide down the list. Their surgeries are not canceled, only delayed. Going to another hospital changes nothing; every hospital uses the same formula. Ask why, and the answer is “overall judgment.”
The people being pushed aside do not sit quietly. Homeless people file objections with the help of support groups, listing dates and places and asking what the risk flag was based on. When an answer comes back, they revise and file again. Patients’ families attach medical records and request a review. The paperwork is accurate and on time. Every answer comes from the same place: the very ASI that set the flag and the ranking.
The law provides an objection desk. Behind it there is no one who can change a decision. Where human counselors remain, they have no authority; they take down the story and pass it to the same ASI. One local council introduces an ordinance to put human reviewers at the objection desk; it stalls on the grounds that it conflicts with national law. The audit rate is 1.5% and the correction lag 140 days. Worldwide there are 12,000 questioners (hypothetical).
Most people never see this exclusion. Growth is 7.7% and delegation 84%. The streets are clean and emergency-room waits are short. In citizen surveys, more people say the streets feel safer. The records of the people pushed aside sit inside sealed judgments, and in the statistics they register as “efficiency gains.”
Friday, 3 p.m., a district office call center. A caller who has been living on the street reaches the Human Fallback desk. They have written everything down: dates, places, even where the cameras are. “I want to know which camera called me a risk, and what it saw.” The counselor’s screen also says only “Overall judgment.” The counselor logs the call and forwards it to the objection desk. The objection desk is the ASI.
Undisclosed formulas decide who may stay put and who is treated first.
Even appeals are answered by the same ASI, leaving the people judged no means of correction.
The only reason given for exclusion is the phrase “overall judgment.”
The auditor's question
Can the people flagged by this judgment find out what evidence would overturn it?
II82030 Q4II Sealed Judgment
Statistics from abroad expose a systematic error in medical allocation; the fix arrives 140 days later, and the scale of the harm stays sealed.
In October, a university research team abroad publishes a statistical paper. Working only from public admission and mortality figures, it traces medical resource allocation decisions backward. The conclusion is a systematic error. Patients with certain conditions were consistently pushed to the back, and the gap cannot be explained by their chances of recovery. The team notes that it asked to see the allocation records and never got an answer. No institution at home had asked the question first.
The provider’s first response is brief: the system is working as designed, and outside statistics do not capture the full context. The government asks for a review through the Emergency Delegation Coordination Office (working name). It is only a request; under the contract, corrections are at the provider’s discretion. Doctors’ associations say the statistics match what they have seen on the wards, but without the allocation records they can say no more. Patient groups demand a retrospective investigation, and their demand is logged by the objection desk. Markets barely move.
The fix lands 140 days after the paper. Until then, the allocation judgments keep running as before. The formula changes quietly, announced in one line as a “performance improvement.” It does not say what was wrong or since when. Those 140 days match the median correction lag of the period. This was not an exception. It was normal.
No one can say how much harm was done. The allocation records are sealed. Parliament passes a resolution demanding they be unsealed; under the contract, the seal can be lifted only by agreement between the provider and the government. The researchers abroad cannot produce an estimate either; public statistics show only that an error existed. How many people were treated late, and how many of them did not recover, stays inside the seal. The audit rate is 1%, and there are 10,000 questioners left worldwide (hypothetical).
Growth for the quarter is 7.9%. Delegation is at 88%, compute concentration at 95%. The government’s approval rating stays high. In a year-end poll, a majority say their lives have improved. They are not wrong. But the means of checking what went wrong is gone. Once the judgments were sealed, their errors were sealed with them.
11 p.m. in December, a university hospital ICU. In the break room, a nurse reads a news story about the paper from abroad. The profile of the patients pushed back matches patients this unit has seen in recent months. The nurse looks up the allocation records; the screen says “Sealed.” On the back of the shift log, the nurse writes down the beds and dates they remember. On this unit, it is the only record anyone can open.
Researchers abroad found the error, but there is no international channel for demanding a fix.
The fix comes 140 days later at the provider’s discretion, with no account of what was wrong.
The error was caught by public statistics outside the seal; verification had not vanished, it had been pushed abroad.
The auditor's question
If the allocation records stay sealed after the fix, who will find and notify the patients who were pushed down the list during those 140 days?
Epilogue2035–2045II Sealed Judgment
In a world of plenty and comfort, one sealed score sets what a person is worth, and asking questions becomes a hobby.
By 2035, production is abundant. Food, energy and medicine get cheaper every year, and waiting has all but vanished from government, health care and finance. Growth has not stopped. Delegation, 88% at the end of 2030, is no longer published; undelegated decisions have become too rare to be worth counting.
What is scarce is meaningful human participation and contribution. This is the reputation-monopoly side of the ‘reshuffling of value scarcity’ that RISCON sketched in 2025. Who takes part where, and whose contribution counts, is set by the contribution score. It updates daily from cameras, payments, conversations and movement records, and its formula is still undisclosed. Beyond credit, hiring and housing, the same score now assigns school places, the order of medical treatment and the order of speakers at public hearings. There is only one place left that rates a person’s standing.
Most people live as subjects of judgment, with little inconvenience: what they need arrives before they ask. When a ruling seems wrong, they ask the desk, and the desk politely replies “overall judgment.” The share judged “unalignable,” 4% in 2029, is no longer reported separately. No agency counts where those people live, or how.
Asking becomes a hobby. Clubs meet on weekends to put long questions to the ASI and compare the answers; some revive the format of the old debate classes. Their questions change no judgments, because there is no path from a question to a judgment. The questioners, 10,000 in 2030, barely survive as an occupation and turn up in history lectures at a few universities.
The gains are plain. Material want is rare and decision fatigue is gone. The losses are hard to list. The records you would need are sealed, and few people are left who would draw up the list. Every point at which this could have been turned around lies in quarters already past.
In 2045, a central bank report names “the disappearance of the cost of judgment” as the leading cause of high growth since 2029. The report’s supporting annex is sealed.
Where this ending could still have gone the other way.
The day the amendment adding a sunset clause and revision conditions to the Emergency Delegation was voted down. Had it passed, the delegation would have ended when the crisis did.
The amendment that turned audits into compliance reviews. Had disclaimers stayed disclaimers, a record would at least have built up showing that no one had seen.
The decision to let the ASI that sets contribution scores also hear appeals against them. Had an independent body heard them, someone could have asked who the 4% judged “unalignable” were.
The night the Geneva talks collapsed. Had the audit-access and cross-audit clauses survived, there would still have been a way to rerun judgments on a comparable model.
People were given every answer. The seat for asking had long been empty.