Products Demo Docs Blog Engineering About Contact Sign in Sign up
Engineering · · Philippe Laporte

Sampling as deterrence: the economics of unpredictable re-execution

You do not have to check every job to make cheating irrational. You have to make sure the executor cannot tell which jobs you will check, and that getting caught costs more than cheating saves. The arithmetic is old, and it is the same arithmetic customs officers and tax auditors use.

The obvious way to verify that an executor ran the model you paid for is to run it again yourself. It works, and it doubles your compute bill. Full cryptographic proof systems, which would let you check the work without repeating it, still cost a large multiple of the inference itself on production-scale models in published, reproducible benchmarks. Newer claims suggest that gap is narrowing, but they have not yet been independently reproduced. Between those two lies sampling: re-run a fraction of the work, chosen in a way the executor cannot predict, and let the economics do the rest.

This piece is about those economics. It uses no construction details, because the argument does not need them. It needs three quantities and one property.

Three quantities

Let c be what it costs an honest executor to run a job correctly.

Let s be the fraction of c the executor saves by cheating on that job: running a smaller model, skipping the computation, returning a cached answer. For model substitution, s might be a large fraction of c. For returning garbage, s is nearly all of c.

Let p be the probability that any given job is independently re-executed and compared. If the verifier re-runs one job in five, p is 0.2.

And let D be what detection costs the executor: the forfeited payment for the job, the contractual penalty, the loss of the contract, the reputational and legal consequences when a counterparty can show, with a signed record, that the executor delivered something other than what it claimed.

The single-job calculation

The expected value of cheating on one job is the saving minus the expected penalty:

expected value of cheating   =   s * c   -   p * D

cheating is irrational when      p * D   >   s * c
equivalently                         D   >   s * c / p

Put numbers in. Suppose the executor could save the entire cost of a job by cheating, so s is 1, and the verifier checks one job in five, so p is 0.2. Honesty is the rational strategy as long as the penalty for getting caught exceeds five times the cost of the job. A contract that forfeits payment on ten jobs for one proven substitution clears that bar with room to spare. A contract that merely withholds payment on the one job does not, and that is worth knowing before you sign it.

Two things fall out of this. First, the sampling rate and the penalty multiply. Halving the sampling rate can be exactly compensated by doubling the penalty, which is why a modest re-execution fraction paired with a serious contractual consequence is enough. Second, the penalty only exists if detection produces something the executor cannot argue with. A verifier's own log is an assertion; a signed record binding the input, the reference output and the executor's output, produced on independent hardware, is evidence. Deterrence rests on the evidence being credible in a dispute.

The repeated-game calculation

Real executors run many jobs, and cheating pays only as a policy, not as a one-off. That is where sampling becomes overwhelming.

If an executor cheats on n jobs and each is checked independently with probability p, the chance of escaping every check is (1 - p) to the power n:

p = 0.2

n =  1   ->  0.8          escapes 4 times in 5
n = 10   ->  0.107        escapes about 1 time in 9
n = 20   ->  0.0115       escapes about 1 time in 87
n = 50   ->  0.0000143    escapes about 1 time in 70,000

An executor who cheats systematically is caught early, almost surely, and every job they cheated on before being caught is now suspect and recoverable under whatever the contract says. An executor who cheats once gains a single job's saving against a one-in-five chance of losing far more. There is no cheating policy that is profitable across a relationship of any length. That is the whole point of sampling: it does not need to catch every bad job; it needs to make bad jobs a losing bet.

The one property that carries everything

None of the above holds if the executor can predict which jobs will be checked.

If the executor knows job 7 will be re-run and job 8 will not, it runs job 7 honestly and cheats on job 8, and the effective sampling rate on the jobs it cheats on is zero. The same applies inside a job: if a large workload is split into pieces and the executor knows which pieces the verifier will look at, it cheats on the rest.

So the selection has to be unpredictable to the executor at the moment it commits to its answer. How a verifier achieves that is a construction detail I will not describe here. The property is what matters, and it is what makes the probability p real rather than nominal.

There is a further subtlety that the customs analogy makes obvious. The selection must be unpredictable, but it must also be uniform, or weighted in a way the executor cannot game. If the verifier only ever re-runs small jobs because they are cheap to check, the executor learns to cheat on large ones. The cost of a check has to be decoupled from the likelihood of a check.

What sampling costs, and what it does not

The overhead of sampled re-execution is roughly p. Check one job in five and the compute bill rises by about a fifth, plus the fixed cost of keeping a verifier on independent infrastructure, which is the part you are actually paying for. Compare that to full re-execution at one hundred percent, or to full cryptographic proofs at a large multiple of the inference on today's reproducible numbers, and sampling is the option that fits inside a normal inference budget now. If proving costs fall as far as some recent claims suggest, that comparison will change, and it should be rerun when they do.

What sampling does not buy is certainty on any individual job. A job that was not selected carries no direct evidence; its assurance is indirect, borrowed from the deterrent effect on the executor's policy. That is a real boundary and it should be stated plainly rather than glossed. For most workloads it is exactly the right trade. For a single high-stakes inference, a decision a regulator will ask about by name, the right sampling rate is one hundred percent, and a well-designed system lets the buyer choose that per job. The cost scales with how much any one answer matters, which is how it should be.

There is also a way to get more detection out of the same budget by committing to the intermediate computation and checking internal consistency rather than re-running end to end. That is research I have written up separately, and it is a different piece of work from the sampling argument here.

The lineage

None of this is new, and it should not be. Gary Becker's 1968 analysis of crime as a rational choice made the same argument about enforcement: the probability of detection and the size of the punishment trade off against each other, and unpredictability is what keeps the probability honest. Random audits, spot inspections and mystery shoppers all work on the same principle. In verifiable computing, sampled re-execution on independent hardware appears in several recent designs, including Proof of Sampling from Hyperbolic and related work on catching dishonest inference providers by re-running selected requests. The economics are the same across all of them.

What to ask a vendor

If a vendor tells you their verification is based on sampling, four questions separate a real design from a nominal one.

Can the executor predict which jobs are checked? If the answer is anything other than a flat no with a mechanism behind it, the sampling rate is fiction.

Can an auditor reproduce the selection after the fact? If not, you are trusting the verifier the way you were trying not to trust the executor.

What is the penalty on detection, and what evidence supports it in a dispute?

Can I raise the rate to one hundred percent for the jobs that matter?


Disclosure: I build sampled independent re-execution at Cyberian Systems. The numbers above are illustrative; they are not our production parameters, and the construction we use to make the selection unpredictable is not described here.

PL
Philippe Laporte
Founder and CEO of Cyberian Systems, building verified AI inference infrastructure for regulated industries.

Read the docs · Try the live demo · RSS