A problem with a hundred generated tests does not need a hundred trips through the new-test dialog. The test script is one program, stored with the problem, that declares all of them: it calls add_test once per test, and running it creates those tests, updates the ones that changed and removes the ones it no longer declares.
The script is written in Starlark, a small language with Python's syntax. It does not generate any data itself. It decides which generator each test calls and with which arguments, and generation then happens the way it does for any other generated test.
The script is edited in the Studio, on the Script view of the Tests tab. Define script, in the menu next to Add test on the problem's Testing tab, opens it there. You need permission to write problems.
Two functions are available.
generator(name, ...) names one of the problem's generators and the arguments to call it with. The arguments are joined with spaces and split again, so generator("gen", "-n", 10) and generator("gen -n 10") are the same command. A name the problem has no generator for fails the run.
add_test(input, answer=None, testset=1, index=None, score=0, example=False, secret=False) declares one test. input and answer are each either a generator command or a string holding the data itself.
Everything else is ordinary Starlark: loops, arithmetic, string formatting, functions.
add_test("3\n1 2 3\n", answer="6\n", testset=1, example=True)
for i in range(1, 11):
add_test(generator("gen", "-n", i * 10, "-seed", i), answer=generator("brute"), testset=2, score=2)
for i in range(20):
add_test(generator("gen -n 200000 -seed %d" % (1000 + i)), testset=3, score=3)Leaving out answer uses the generator named solution, which is the convention for computing answers with the author's own program. Leaving out index gives the test the lowest index in its testset that no hand-made test and no earlier add_test has taken. testset is the testset's number as the Testing tab shows it; the script never creates testsets, so add those first.
Preview runs the script without saving anything and lists what it would do to the problem's tests in a Script Output tab: Added, Changed, Removed and Unchanged, one row each, with the generator command that produces it. Anything the script printed appears above the table.
Read the Removed rows before running. A test the script produced last time and does not produce now is deleted.
Run saves the script and applies exactly what the preview showed. All of it lands as a single change to the problem, so one generation task covers every test the run touched — watch it on the problem's Activity tab along with the space's other background tasks.
An error stops the run before anything is written, and the message names the line it happened on. A run that would take the problem over its test limit is refused.
Running the script matches what it declares now against what it produced last time, by testset and index. A test at an index it produced before keeps its identity and is updated in place; an index it did not use before becomes a new test; an index it no longer declares is removed.
Tests you created by hand are never touched, and their indexes are not available to the script. Asking for one of them explicitly is an error, and leaving the index out skips over them.
The script does not lock the tests it produced. You can edit one from the test list like any other, and the change holds until the next run, which reconciles that test back to what the script declares.
Running an empty script removes the script and every test it produced.
The script cannot read the clock, generate random numbers or reach the network, so running it twice in a row produces the same tests both times.
Randomness belongs in the generator, which draws it from the arguments the script passes. Two tests calling one generator with identical arguments are two copies of the same test, so vary something — a seed argument is the usual choice. See Generate tests for why a generator has to be deterministic in the first place.
Adding a validator is worth it once a script is producing tests in bulk: it checks that every generated input actually satisfies the problem's constraints, which is the failure a large family of generated tests is most likely to hide.