Stress testing

A stress test hunts for the test you didn't think of. It runs one of your generators over and over with random arguments, and on every input it produces it compares your solutions against a reference one. The moment a solution does something its declared type does not allow, the run stops and hands you the exact generator command that produced the input.

This is the tool for the case where you know a solution is wrong — or suspect it might be — but have no idea which input exposes it. Write a slow brute force you trust, point the stress at it, and let it find the counterexample.

You start a stress from the Stress testing row on the problem's Testing tab. You need permission to write problems.

Before you start

You need three things in place:

  • a generator that takes its parameters from the command line, so a stress can vary them;

  • a reference solution of type Correct — its output is the answer everything else is measured against;

  • at least one other solution to compare against it.

A validator is not required, but with one in place a generator that produces an input outside the problem's constraints is reported as a generator problem instead of being blamed on a solution. The problem's checker is used to compare outputs, so problems with several acceptable answers work correctly.

Running a stress

  1. Open the problem's Testing tab and select Stress testing.

  2. Pick the Generator.

  3. Write the Arguments — the generator's command line, with the parts you want randomised written as ranges. See below.

  4. Pick the Reference solution. Only solutions of type Correct are offered; the first one is already selected.

  5. Choose the Compared solutions, or leave the field empty to compare every solution the problem has.

  6. Adjust Runs, Time budget and Time limit per run if you want, then Run. Left empty, the time limit per run is the problem's own.

The run opens on its own page and fills in as results arrive.

Writing the arguments

Arguments are the command line your generator receives. Write a [a..b] token wherever you want a random integer:

-n [1..100] -max [1..1000000]

On each run, [1..100] becomes some number between 1 and 100 inclusive. A random seed argument is added on top of whatever you write, so a fixed command like 1000 still produces a different test every time — testlib-style generators derive their randomness from the arguments they are given, and the seed is what makes each run differ.

Ranges are the only pattern accepted; anything else in the arguments is passed through as written.

Start wide. When a counterexample turns up, narrow the ranges around it and run again — that is how you get from "it breaks somewhere under 100000" to a small test you can read.

Reading the result

The page lists every run with the arguments it used and its verdict:

Verdict

Meaning

Passed

Every solution behaved the way its type promises.

Counterexample

One did not. This is the one you came for.

Invalid input

The validator rejected the generated input. The generator is at fault, not a solution.

Broken

The generator, the reference solution or the checker did not finish.

When a counterexample is found it gets a card of its own at the top of the page, with the generated input, the reference solution's answer, and a line per compared solution showing its verdict and whether that verdict is one its type allows. Selecting a solution opens its output.

A verdict is judged per test, against the solution's declared type. A solution marked Timeout is allowed to pass a small random input — only a verdict its type does not allow, such as a wrong answer from a solution that is merely supposed to be slow, makes the run a counterexample. A solution marked Incorrect never produces one, because it is expected to fail by any means. This is deliberately looser than Submit all on the Solutions tab, which judges a solution across the whole testset.

Keeping a counterexample

Selecting any run offers Add as test. It opens the new-test form with the generator and that run's exact arguments already filled in, so the test stores the command rather than a file — the input is regenerated the same way as every other generated test. Pick the testset it belongs in and save.

This is the only thing that survives the run. A stress is kept for an hour and is then gone.

While it runs

A stress occupies a judge for as long as it runs, so it is bounded: at most 500 runs and 10 minutes of wall time, and one stress per problem at a time. Starting a second while one is in flight is refused — cancel the first if you want to change something.

Cancel stress stops results being collected. Work already handed to the judge finishes on its own, so the run may keep costing judge time until its budget runs out; what it produces after cancellation is dropped.

By default a run stops at the first counterexample, which is usually what you want. Keep going after the first counterexample spends the whole budget instead, which is useful when you want to see how often a solution fails rather than just whether it does.

When the run fails outright

If a program does not compile, the run ends immediately with the compiler's output on the page and no runs at all. A generator or reference solution that crashes at run time ends the run as Broken instead — the stress cannot tell you anything about your solutions when the thing producing the answers is itself failing.