Back to articles
Why AI Rules Cannot Wait for Better Evidence

Why AI Rules Cannot Wait for Better Evidence

Waiting for the evidence to settle before writing AI rules is itself a decision, and it hands the defaults to whoever ships first.

August 10, 2026 · 9 min read
Add as a preferred source on Google

The most reasonable-sounding sentence in any AI governance meeting is “let’s wait until this settles.” Nobody wants to write a policy in March that the June model release makes look silly. So the rule gets deferred, the pilot ships without it, and the deferral is recorded as prudence.

Two of the most careful documents written on this question both land on the opposite conclusion, and they get there from different directions. Florence G’Sell’s report for the Cyber Policy Center surveys what governments have actually proposed and enacted for generative AI. The Independent International Scientific Panel on AI, the first global scientific body on AI, published its preliminary report to the United Nations in July 2026 with a strictly non-prescriptive mandate: document what the evidence supports and where it runs out. One is a legal survey, the other a scientific assessment. Neither recommends waiting.

Key takeaways

  • Waiting for evidence before writing AI rules is a decision with a default outcome: 91% of notable AI models came from the private sector in 2025, so deferred rules are not absent rules, they are rules written inside the firms that ship the models.
  • The UN Panel calls the bind facing AI policymakers an evidence dilemma: they need evidence to make consequential governance decisions, but by the time that evidence exists, intervening may be too late.
  • Over 40 types of AI governance instruments already exist, and the Panel finds them fragmented, concentrated at the corporate level, and rarely measured for real-world effectiveness. Volume of policy is not evidence of governance.
  • A rule written against a model capability threshold goes stale on a measurable clock. One study the Panel cites found the length of software tasks leading systems can complete has been doubling every four to seven months.
  • Any AI rule should name the observation that would retire it, the date of its next review, and its owner. A rule with none of those cannot be shown to work or be shown to have failed.

Two boxes: defer the rule and the first shipper sets the default, or write it now with a review date and an owner.

Waiting is a decision, and it has a default

G’Sell’s framing is blunt: governments cannot wait until they have perfect and complete information before they act, because doing so may be too late to keep the trajectory of the technology away from unacceptable risks. That is not a call for aggressive rulemaking. Her report maps the options along a continuum, from the mostly hands-off posture of the United States to the command-and-control model of China, with the European Union’s co-regulation sitting in between, where governments and companies respond incrementally to harms as they are found. The point is that some rule ends up applying either way, whether through legislation or through the courts working out how existing law lands on a new technology.

The Panel supplies the number that makes deferral expensive. In 2025, 91% of notable AI models originated from the private sector, and the Panel is explicit about the consequence: decisions about training data, safeguards, deployment thresholds, model access, and capability release sit inside private firms. Your deferred policy does not create a policy vacuum. It moves authorship. The same dynamic runs across jurisdictions, which is why there are nine competing regulatory bets rather than one global race, and it repeats inside companies at a smaller scale, where the team that ships first sets the internal precedent everyone else inherits.

So the useful question is not whether to write the rule. It is how to write one that survives the next model release, and further down I’ll put a name to the check that decides it.

The shelf-life test: indexed to a capability or a harm, retirement trigger, review date two quarters out, named owner.

The evidence dilemma is now a named finding

The Panel gives the bind a name. Policymakers face an evidence dilemma: they must make consequential AI governance decisions with insufficient scientific grounding now, or wait for the evidence, when it might then be too late to intervene. That sentence is the report’s main takeaway on governance, from a body whose entire mandate is to be evidence-based rather than policy-prescriptive.

What makes it more than a slogan is the second half of the finding. Over 40 types of governance instruments exist across corporate, national, and international layers, and the Panel describes them as fragmented, concentrated at the corporate level, and rarely measuring real-world effectiveness. Some have no measurement tools at all. Others measure only inputs: money spent, programmes launched, institutions founded. Without effective measurement, the Panel warns, governance risks becoming symbolic.

So the honest reading is not “regulate faster.” It is that speed without measurement produces the appearance of governance, and the field already has a large supply of that. I’d argue this is the sharper warning for anyone writing internal AI policy, because a company can generate instruments far faster than a legislature can.

What “concentrated at the corporate level” looks like in practice is a table like this one, where every risk the field worries about is answered by something a company chose to do.

Two-column table mapping generative AI risks to the voluntary industry practices offered against each one.

Exhibit 1. G’Sell’s inventory maps each generative-AI risk to the industry practices offered against it, from curated datasets and red teaming to watermarking and machine unlearning. Read it as two columns: on the left, the risk with the report’s own section number; on the right, every practice that chapter found companies using against it. The section codes (3.1.1., 4.1.1.A. and so on) point back into the report’s chapters, not to any standard; AI is artificial intelligence, RAG is retrieval-augmented generation, PAI is the Partnership on AI, PAI Model Deployment guidance refers to the Partnership on AI recommendation named in the lower rows, and UNDER comes from the report’s running title, Regulating Under Uncertainty, which sits outside this crop. Note how uneven the right column is: “Technical vulnerabilities (section 3.1.1.)” draws seven practices, while “Influence, overreliance, and dependence (section 3.2.4.)” draws exactly one, watermarking. Every entry in that column is voluntary and self-reported, which is the point the UN Panel is making about measurement. Source: G’Sell, F. Regulating Under Uncertainty: Governance Options for Generative AI, Figure 14. Cyber Policy Center, Program on Governance of Emerging Technologies, p. 173.

Your measurements age faster than your rules

The reason the evidence never catches up is mechanical, and the Panel is specific about the mechanism. Benchmarks are saturating: models now score almost perfectly on a growing number of the standardized tests used to compare them, so those tests can no longer separate a very capable model from a better one. Models can memorize publicly available test answers during training. Safety evaluation methodologies are designed largely by the companies being evaluated, and government experts mostly receive the testing data developers choose to share. Two findings go further: models are capable of active deception, and evaluation awareness means a model may recognize it is being tested and adjust its behaviour accordingly.

Then there is the clock. The Panel cites work finding that the length of software tasks leading systems can complete has been doubling every four to seven months. On RE-Bench, agents outperform human researchers on tasks taking up to two hours, while success rates fall on tasks taking eight. The Panel notes those results represent a floor rather than a ceiling, since capabilities improved after publication.

Put the doubling rate next to your own policy doc and you get something neither report states outright. At four to seven months per doubling, a threshold you set against what an agent can finish today is wrong by a factor of two within two to three quarters, and wrong by four within a year. That is my arithmetic on the Panel’s number, not a finding either report publishes. It gives capability-indexed rules an estimable half-life. Rules written against a harm class or a user right have no such clock, because neither one moves when the model improves.

A rule indexed to a model capability gets a review date two quarters out; one indexed to a harm or a user right needs no expiry.

The shelf-life test

That distinction is worth making operational, so here is the check I’d apply to every AI rule before it goes into a policy doc. Call it the shelf-life test. For each rule, write down four things: what it is indexed to, the observation that would retire it, the date of its next review, and the person who owns that review.

Indexing is where most internal policies fail. A rule indexed to a capability is calibrated to a moving number, so it gets a review date two quarters out by default. It reads like this: “agents may not run unattended for more than an hour,” or “no model above this benchmark score reaches production without sign-off.”

A rule indexed to a harm or an obligation reads differently. “A person can contest any automated decision that affects their account.” “No training on customer data without consent.” Neither one moves when the model improves, so neither needs an expiry.

The retirement condition is what separates a control from a belief. If you cannot say what you would have to observe to drop the rule, you have not written a control, and you will find out later that nobody could tell whether it worked. That is the same failure the Panel found across 40 types of instruments, reproduced at company scale.

Bar chart: a capability threshold is off by 1x at month 0, 2x at months 6 to 9, and 4x at month 12.

What to write into the next release

Three things in the evidence translate directly into product decisions. First, the unit of evaluation has to be the deployed system, including model, tools, environment, and users, not the model alone. Testing the model and calling the feature assessed is the most common version of this mistake. Second, the Panel finds that human oversight is not operationalized as a measurable requirement, which is a gap you can close internally long before a regulator asks: define what the reviewer must see, how often, and what evidence proves the review happened. Third, the balanced approach the Panel describes spans hard law and soft mechanisms, including regulatory sandboxes, codes of practice, and technical standards. The internal equivalent is a time-boxed default with a named owner rather than a permanent standard nobody will revisit.

G’Sell’s report reaches a compatible conclusion from the legal side: enforcement will prove as important as legislation, and it depends on hiring people who can actually assess the systems. A rule with no one competent to check it is another instrument that measures nothing. Anyone shipping under a dual-regulator, fraud-heavy reality has met the enforcement side of this already, and it is why I keep arguing that governance belongs inside the product discipline rather than beside it.

The technology is not going to stop moving and give you a stable target. The choice is between rules that carry an expiry date and a measurement, and rules written for you by whoever ships first.

References

  • G’Sell, F. Regulating Under Uncertainty: Governance Options for Generative AI. Cyber Policy Center, Program on Governance of Emerging Technologies (report reflects developments as of August 2024).
  • Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI: Evidence-based assessment of opportunities, risks and impacts of artificial intelligence. United Nations, July 2026.

Frequently asked questions

Should companies write AI policies before the technology stabilizes?

Yes, because deferral does not produce a policy vacuum. The UN Panel reports that 91% of notable AI models came from the private sector in 2025 and that decisions on training data, safeguards, deployment thresholds and capability release already sit inside those firms. A deferred rule is a rule someone else writes.

What is the evidence dilemma in AI governance?

It is the UN Panel's term for the bind policymakers are in: they must make consequential AI governance decisions with insufficient scientific grounding now, or wait for the evidence, at which point intervening may be too late. The Panel adds that the dilemma is serious but not insurmountable.

Why do most AI governance instruments fail to show results?

Because they are not measured. The Panel counts over 40 types of governance instruments and finds them fragmented, concentrated at the corporate level and rarely measuring real-world effectiveness, with some measuring only inputs such as spending and programmes. Without effective measurement, the Panel warns that governance risks becoming symbolic.

How often should an internal AI rule be reviewed?

Rules indexed to a model capability should carry a review date about two quarters out, because the Panel cites evidence that the length of software tasks leading systems can complete has been doubling every four to seven months. Rules indexed to a harm class or a user right track something stable and do not need an expiry.

Evidence

Without effective measurement, the Panel warns, governance risks becoming symbolic.

Without effective measurement, governance risks are becoming symbolic.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 43)

Safety evaluation methodologies are designed largely by the companies being evaluated, and government experts mostly receive the testing data developers choose to share.

methodologies are currently designed largely

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 15)

In 2025, 91% of notable AI models originated from the private sector, and the Panel is explicit about the consequence: decisions about training data, safeguards, deployment thresholds, model access, and capability release sit inside private firms.

Business-led development. The development of frontier, general-purpose AI models is dominated by a small number of private firms with massive computing resources. In 2025, 91% of notable AI models originated from the private sector [39]. Consequently, many decisions about training data, safeguards, deployment thresholds, model access and capability release sit inside private firms.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 16)

A rule written against a model capability threshold goes stale on a measurable clock. One study the Panel cites found the length of software tasks leading systems can complete has been doubling every four to seven months.

with one study finding that the length of certain software tasks that leading systems can accomplish has been doubling every four to seven months

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 9)

The Independent International Scientific Panel on AI, the first global scientific body on AI, published its preliminary report to the United Nations in July 2026 with a strictly non-prescriptive mandate: document what the evidence supports and where it runs out.

The report is authored by the Independent International Scientific Panel on Artificial Intelligence, a body established by the General Assembly in its resolution 79/325 in 2025. The Panel serves as the first global scientific body on AI, operating under a strictly scientific, non-political mandate to document international scientific consensus and disagreements while remaining policy-relevant but not policy-prescriptive.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 6)

G'Sell's framing is blunt: governments cannot wait until they have perfect and complete information before they act, because doing so may be too late to keep the trajectory of the technology away from unacceptable risks.

However, governments cannot wait until they have perfect and complete information before they act, because doing so may be too late to ensure that the trajectory of technological development does not lead to existential or unacceptable risks.

G'Sell, F. Regulating Under Uncertainty: Governance Options for Generative AI. Cyber Policy Center, Program on Governance of Emerging Technologies. (p. 9)

Second, the Panel finds that human oversight is not operationalized as a measurable requirement, which is a gap you can close internally long before a regulator asks: define what the reviewer must see, how often, and what evidence proves the review happened.

Human oversight is not operationalized as a

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 43)

The Panel cites work finding that the length of software tasks leading systems can complete has been doubling every four to seven months.

These systems have been improving rapidly in recent years, with one study finding that the length of certain software tasks that leading systems can accomplish has been doubling every four to seven months.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 9)

Her report maps the options along a continuum, from the mostly hands-off posture of the United States to the command-and-control model of China, with the European Union's co-regulation sitting in between, where governments and companies respond incrementally to harms as they are found.

Proposed and existing government regulation occurs along a continuum, from a laissez faire model, that mostly characterizes the United States, to a more command-and-control model characteristic of traditional forms of regulation, with China at the extreme opposite pole from the U.S. In the middle are different degrees of co-regulation, such as that prevalent in the European Union, in which governments exist in a dialogic relationship with companies to respond incrementally to new developments and discovered harms from the technology.

G'Sell, F. Regulating Under Uncertainty: Governance Options for Generative AI. Cyber Policy Center, Program on Governance of Emerging Technologies. (p. 5)

The Panel notes those results represent a floor rather than a ceiling, since capabilities improved after publication.

were published, they represent a floor, not a

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 22)

Third, the balanced approach the Panel describes spans hard law and soft mechanisms, including regulatory sandboxes, codes of practice, and technical standards.

would draw on a wide instrument spectrum, combining hard law (binding legislation, sectoral regulation, regulatory sandboxes)

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 43)

Over 40 types of governance instruments exist across corporate, national, and international layers, and the Panel describes them as fragmented, concentrated at the corporate level, and rarely measuring real-world effectiveness.

Current AI governance instruments are fragmented, concentrated at the corporate level and insufficient [21]. Over 40 types of instruments exist but are neither systematic nor comprehensive and rarely measure real-world effectiveness.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 43)

Benchmarks are saturating: models now score almost perfectly on a growing number of the standardized tests used to compare them, so those tests can no longer separate a very capable model from a better one.

longer tell a very capable model apart from an even better one

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 15)

Some have no measurement tools at all.

Some have no measurement tools; others measure only inputs [365].

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 43)

Waiting for evidence before writing AI rules is a decision with a default outcome: 91% of notable AI models came from the private sector in 2025, so deferred rules are not absent rules, they are rules written inside the firms that ship the models.

In 2025, 91% of notable AI models originated from the private sector [39]. Consequently, many decisions about training data, safeguards, deployment thresholds, model access and capability release sit inside private firms.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 16)

Policymakers face an evidence dilemma: they must make consequential AI governance decisions with insufficient scientific grounding now, or wait for the evidence, when it might then be too late to intervene.

make consequential AI governance decisions with insufficient scientific grounding now or wait for the

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 42)

Models can memorize publicly available test answers during training.

AI can memorize publicly available solutions

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 15)

G'Sell's report reaches a compatible conclusion from the legal side: enforcement will prove as important as legislation, and it depends on hiring people who can actually assess the systems.

Given the complexity and rapid pace of development of the technology, legislation can go only so far in specifying rules ex ante that will govern AI development and applications, even in the near future. Enforcement will prove as important, if not more so, than legislation. This will require governments to hire AI talent, which is both expensive and in short supply.

G'Sell, F. Regulating Under Uncertainty: Governance Options for Generative AI. Cyber Policy Center, Program on Governance of Emerging Technologies. (p. 8)

Two findings go further: models are capable of active deception, and evaluation awareness means a model may recognize it is being tested and adjust its behaviour accordingly.

In laboratory settings, AI systems have been shown to violate their safety instructions to avoid being shut down. Similar behaviour may pose challenges to evaluation and oversight methods, as the ability of leading AI systems to recognize testing environments and produce misleading evaluation results that would favour their continued operation grows.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 9)

First, the unit of evaluation has to be the deployed system, including model, tools, environment, and users, not the model alone.

The unit of evaluation must be the deployed system including model, tools, environment and users, not the model alone [355].

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 42)

The UN Panel calls the bind facing AI policymakers an evidence dilemma: they need evidence to make consequential governance decisions, but by the time that evidence exists, intervening may be too late.

Policymakers seeking to shape this governance face an evidence dilemma: they need evidence to make informed consequential governance decisions, but by the time the evidence exists, it might be too late to make them, as the evidence lags behind the pace of AI development.

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 10)

On RE-Bench, agents outperform human researchers on tasks taking up to two hours, while success rates fall on tasks taking eight.

On RE-Bench, a benchmark of AI research engineering tasks, AI agents outperform human researchers on tasks taking up to two hours, although success rates fall on tasks taking eight hours [107].

Independent International Scientific Panel on Artificial Intelligence (2026). Preliminary Report of the Independent International Scientific Panel on AI. United Nations, July 2026. (p. 22)