What does Dr. Jonah Tebaa on the 24-Hour Standard for AI Failure Disclosure mean in practice?
Dr. Jonah Tebaa proposes a four-window disclosure clock and pre-approved notification threshold to govern AI failures within twenty-four hours. His framework structures response time instead of rushing public statements: detecting and containing damage in hours one to three, quantifying the blast radius in hours three to six, determining who meets notification thresholds in hours six to twelve, and conducting individual notifications alongside drafting public statements across hours twelve to twenty-four.
Between 11 p.m. and 6 a.m., an overnight batch pass from a regional retail bank's fraud-detection model froze 1,340 customer accounts it had flagged as high-risk. By the time anyone with authority looked closely at what the system had done, most of that night's damage was already locked in. In a recent piece, Dr. Jonah Tebaa uses this composite scenario — assembled from patterns across several fraud and risk deployments, not a real institution or client — to argue that the ethics of an AI failure are settled long before the failure itself makes headlines.
His claim is specific and, once you see the arithmetic, hard to dismiss: what determines whether an AI incident is handled responsibly is not the tone of the apology or the speed of the press release. It is whether a disclosure clock and a notification threshold already existed, in writing, before the system made its first mistake.
An alert that fired and was never seen
Three hours into the batch run in his example, the freeze-rate monitor crossed its own alert threshold. Four hundred accounts had already been locked, a rate that would normally trigger review. But a suppression rule, written at some earlier point to treat weekend-night volume as expected variance, caught the alert automatically. No human ever saw it.
Dr. Tebaa treats this as the pivotal moment in the whole incident — not the freezing itself, but the quiet second decision that determined nobody would look at it for six more hours. He argues that monitoring thresholds function as policy decisions dressed up as technical settings, and that a threshold a scheduling rule can silently override is, in effect, a threshold nobody actually owns.
The measurement that has to happen before any statement
Call-centre volume spiked sixfold at hour nine, prompting a supervisor to escalate manually — the human intervention that should have happened six hours earlier. By then the batch had finished: 1,340 accounts frozen in total. Manual review found that 1,193 of them, 89 percent, were false positives, legitimate customers locked out of their own money with an average frozen balance of $2,300, some of it rent or payroll. Only 147 accounts, 11 percent, involved genuine fraud.
The point Dr. Tebaa draws from this sequence is a discipline, not a detail: an organization cannot ethically communicate what it has not yet measured. A statement issued before the blast radius is known is guesswork wearing the costume of transparency. In his framework, the first hours after any AI failure exist to quantify harm, not to manage perception.
A clock with four jobs, not one deadline
The framework he proposes replaces the instinct to "respond quickly" with four distinct windows, each with its own job, agreed before any incident occurs: detect and contain in the first three hours; quantify the blast radius in hours three to six; decide, against a pre-set policy, who crosses the notification threshold in hours six to twelve; and only then, in hours twelve to twenty-four, begin individual notification and draft any public statement.
Crucially, none of the first three windows involve public messaging at all. The apology, if warranted, belongs at the end of the clock — once the organization actually knows what happened and to whom.
Why the threshold has to be set in advance
The most exposed moment in Dr. Tebaa's account is not the failure itself but the decision at hour six about who counts as "affected enough" to notify. In his example, that decision was effectively made twice: once by accident at hour three, when a suppression rule decided the breach wasn't worth flagging, and once for real at hour six, under a call centre in chaos, by people improvising a threshold they had never rehearsed.
His argument is that a notification threshold decided in the middle of a live incident is not a policy — it's whoever is most persuasive in the room that hour. The alternative he recommends is a written, pre-approved threshold specifying exactly what triggers individual notification and what triggers a public statement, by count of people affected, dollar exposure, and harm category. He is careful to note that actual disclosure obligations vary by jurisdiction and regulator, and that his framework is a decision-making structure, not a substitute for legal counsel.
The cost of six hours with no one at the wheel
The number Dr. Tebaa keeps returning to is 940 — the accounts frozen between hour three and hour nine, while a threshold breach sat unseen. Four hundred were frozen when the alert fired; 940 more followed before a human intervened; 400 plus 940 equals the full 1,340. Individual notification eventually went out at hour fourteen, and a public statement followed on day three, at hour seventy-two — both, in his telling, reasonable once the clock finally started running.
Those 940 accounts are what he calls the direct cost of an undecided clock: not bad luck, but the absence of a pre-committed answer to a question the organization should have already asked itself — what does this alert mean, and who needs to see it? It is a framework built less around better software than around deciding, ahead of time, exactly who owns that question.