Checkers with testlib.h

When this page applies. This is the testlib version of Write a checker, which uses eolymp.h and is the default for a new checker. Use this page only when:

  • the checker being changed already uses testlib — it includes testlib.h or calls registerTestlibCmd (read it first with get_checker). Change it in testlib; do not port it to eolymp.h unless the requester asks for exactly that; or

  • the requester explicitly asks for a testlib checker.

The rule, and why, is in eolymp.h reference.

1. What a checker is

A program that decides whether a contestant's output is correct for one test, and exits witha verdict. It is given three filenames as argv, exposed by testlib as three streams:

stream

is

trust

inf

the test input

trusted — the validator already vetted it

ouf

the contestant's output

hostile — arbitrary bytes, assume nothing

ans

the jury's answer

verify it too — see §4

Write a PROGRAM checker only when correctness genuinely needs logic — several validanswers, a constructed object to verify, or partial credit. Reach for a built-in first:

  • LINES — byte-for-byte line comparison. Trailing spaces and tabs on each line, andempty lines at the end of the file, are ignored; everything else must match exactly.

  • TOKENS — compares word by word or number by number, deciding which by what the answerfile holds. Numbers are compared with a relative tolerance, (and when ), inlong double. Words are compared case-insensitively unless case_sensitive is set.Slower than LINES, so avoid it on very large outputs.

  • QUERY_RESULTS — for per-query output.

A statement promising "any case is accepted" is only true under TOKENS withcase_sensitive off. Under LINES, or under a custom checker that does not fold case, thatsentence is a lie that costs someone a submission.

LEGACY_PROGRAM takes its arguments in a different order. A PROGRAM checker is called<input> <output> <answer> — testlib's order, which registerTestlibCmd expects. ALEGACY_PROGRAM checker is called <input> <answer> <output>. Choosing the wrong typesilently swaps the contestant's output with the jury's answer, and the checker still runs.

There is no need to validate the input in a checker — inf already passed the validator,so read it without readSpace/readEoln.


2. The skeleton

#include "testlib.h"
using namespace std;

int main(int argc, char* argv[]) {
    registerTestlibCmd(argc, argv);

    int n = inf.readInt();                                       // trusted
    int jans = ans.readInt();                                    // jury
    int pans = ouf.readInt(-2000, 2000, "sum of numbers");       // BOUNDED + NAMED

    if (pans != jans)
        quitf(_wa, "expected %d, found %d", jans, pans);
    quitf(_ok, "the sum is %d", jans);
}

That shape is only adequate when the answer is a single scalar. For anything composite,use §4 instead.


Runtime: store it as cpp:20-gnu14. Runtime ids are language:version-toolchain;a bare cpp17 or cpp is rejected — see testlib reference.


3. Verdicts

call

meaning

quitf(_ok, ...)

correct, and optimal if applicable

quitf(_wa, ...)

wrong — contestant's fault

quitf(_pe, ...)

malformed output, format violated

quitf(_fail, ...)

the jury is wrong or the checker broke

quitp(score, ...)

partial credit

The full TResult set from testlib.h: _ok = 0, _wa = 1, _pe = 2, _fail = 3,_dirt = 4, _points = 5, _unexpected_eof = 8, _partially = 16.

How the judge reads the exit code: 0 accepted · 1 or 2 wrong answer / presentationerror · 7 partial points, and the judge then scans the log for the token points followedby a number — exactly what quitp emits · anything else is a system failure.

Helpers worth knowing:

quitif(cond, _wa, "…", ...);   // quitf(...) when cond is true
ensuref(cond, "…", ...);       // assert; fires _fail — for checker-internal invariants

_fail is not a contestant verdict

_fail means the problem is broken — a checker bug, an incorrect jury answer, or acontestant who found something better than the jury. It demands jury investigation. Neveruse it for anything a contestant can legitimately trigger, and never return _wa when thejury answer is impossible: that hides a broken problem behind contestant failures.

_pe vs _wa

_pe is for output that cannot be parsed; _wa for output that parses but is wrong. OnCodeforces (and Eolymp) _pe is largely folded into _wa, so the distinction is mostlydiagnostic. Note that testlib's bounded readers already emit _pe on garbage, so youusually get it for free — another reason bounds matter (§5).

_pc vs quitp

_pc(score) takes an integer 0–200, not 0–100: 0 is no points, 200 is maximum. It isalso historically quirky — a documented +16 offset makes quitf(_pc(50), …) not mean whatit looks like. Use quitp(score, "…") instead. See §7 for what Eolymp does with thatnumber, which is not what CMS does.


4. The readAns paradigm — use it whenever the answer is composite

This is the single most important structural rule for a non-trivial checker, and it is thecore recommendation of the reference blog.

Write one function that reads an answer from a stream, validates it, and returns itsvalue. Call it on ans first, then on ouf.

#include "testlib.h"
#include <map>
#include <vector>
using namespace std;

map<pair<int,int>, int> edges;
int n, m, s, t;

// Reads an answer from `stream`, checks it really is a simple s–t path,
// and returns its weight.
int readAns(InStream& stream) {
    int value = 0;
    vector<int> path;
    vector<bool> used(n, false);

    int len = stream.readInt(2, n, "number of vertices");
    for (int i = 0; i < len; i++) {
        int v = stream.readInt(1, n, format("path[%d]", i + 1).c_str());
        if (used[v - 1])
            stream.quitf(_wa, "vertex %d was used twice", v);
        used[v - 1] = true;
        path.push_back(v);
    }
    if (path.front() != s)
        stream.quitf(_wa, "path doesn't start at s: expected %d, found %d", s, path.front());
    if (path.back() != t)
        stream.quitf(_wa, "path doesn't finish at t: expected %d, found %d", t, path.back());
    for (int i = 0; i + 1 < len; i++) {
        if (!edges.count(make_pair(path[i], path[i + 1])))
            stream.quitf(_wa, "there is no edge (%d, %d)", path[i], path[i + 1]);
        value += edges[make_pair(path[i], path[i + 1])];
    }
    return value;
}

int main(int argc, char* argv[]) {
    registerTestlibCmd(argc, argv);
    n = inf.readInt(); m = inf.readInt();
    for (int i = 0; i < m; i++) {
        int a = inf.readInt(), b = inf.readInt(), w = inf.readInt();
        edges[make_pair(a, b)] = edges[make_pair(b, a)] = w;
    }
    s = inf.readInt(); t = inf.readInt();

    int jans = readAns(ans);        // JURY FIRST — see below
    int pans = readAns(ouf);

    if (jans > pans)
        quitf(_wa,   "jury has the better answer: jans = %d, pans = %d", jans, pans);
    else if (jans == pans)
        quitf(_ok,   "answer = %d", pans);
    else
        quitf(_fail, ":( participant has the better answer: jans = %d, pans = %d", jans, pans);
}

The mechanism: stream.quitf

stream.quitf(...) rewrites the verdict according to which stream it was called on. Onouf it behaves exactly like quitf. On ans any verdict becomes _fail, because amalformed jury answer is a jury bug, not a contestant error.

That one detail is what makes a single function serve both roles. Plain quitf insidereadAns would report a broken jury answer as the contestant's wrong answer — the exactbug the paradigm exists to prevent.

Call readAns(ans) before readAns(ouf)

Order matters and is not cosmetic. Validating the jury first means a broken problemsurfaces as _fail immediately. Reverse the order and a contestant with a wrong answerquits the checker at _wa before the jury answer is ever parsed — so the broken juryanswer stays hidden for as long as no correct submission arrives. Jury first, always.

Why this is the default, not an optimisation

The alternative — parsing ouf inline and reading the jury with a bare ans.readInt() —has two defects:

  1. It trusts the jury absolutely. An unbounded, unstructured read of ans verifiesnothing. If the model solution is wrong, every correct contestant gets _wa.

  2. It duplicates the extraction logic, or skips the jury side entirely — which is worse.For a composite answer the extraction is the complicated part, and having it twice meansfixing it twice when the output format changes.

The paradigm applies to any composite answer, including NO / (YES + certificate)problems: readAns returns whether the stream said yes, having verified the certificate.


5. Never trust ouf

This is the most common real defect in checkers, and the reference blog leads its "commonmistakes" with it.

int k = ouf.readInt();                 // BAD — accepts 2147483647, or -5
vector<int> lst(k);                    // …allocates 8 GB, or constructs with a negative size

int pos = ouf.readInt();
int x = A[pos];                        // BAD — "100% place for runtime-error"
int k   = ouf.readInt(0, n, "k");                       // GOOD
int pos = ouf.readInt(0, (int)A.size() - 1, "pos");     // GOOD

A checker that crashes is reported by the judge as a system error, not a wrong answer,and can take down a whole rejudge. Rules:

  • Every ouf.read* producing a count, index or length is bounded by something derivedfrom inf.

  • Every value is range-checked before it indexes anything.

  • Cap total reads; no unbounded while (!ouf.seekEof()) over untrusted output.

  • Always pass the name argument. Without it the failure reads "integer doesn't belong torange [23, 45]" with no indication of which integer. Contestants often see these messages.

Trailing content: ouf.skipBlanks(); if (!ouf.seekEof()) quitf(_wa, "extra output");.Use the tolerant readers (skipBlanks, seekEof, seekEoln) in a checker — be lenientabout a contestant's trailing whitespace — and the strict ones (readSpace, readEoln,readEof) in a validator, where the jury's own data must be exact.


6. Problems with many valid answers

Most checkers never read ans at all. For problems where the answer is verifiable from theinput alone — "output any valid colouring", "output any shortest path" —the checker recomputes validity and never needs the jury.

That is still readAns shaped. Write the verifier as a function over a stream even when youonly call it once, so the day the problem gains an optimality criterion you addreadAns(ans) above the existing call rather than restructuring:

long long readAns(InStream& stream) { /* validate, return the objective value */ }

int main(int argc, char* argv[]) {
    registerTestlibCmd(argc, argv);
    /* read inf */
    long long jans = readAns(ans);      // jury first
    long long pans = readAns(ouf);
    if (pans < jans) quitf(_wa,   "jury %lld beats participant %lld", jans, pans);
    if (pans > jans) quitf(_fail, "participant %lld beats jury %lld", pans, jans);
    quitf(_ok, "answer = %lld", pans);
}

For optimisation problems always check both directions — the answer is feasible, andits value matches what the contestant claims and what the jury achieved. A contestant whobeats the jury is _fail; accepting it hides a broken model solution for years.


7. Partial scoring

quitp(points, "scored %d of %d", points, maxPoints);

Three platform behaviours silently corrupt scoring if you get them wrong.

7.1 Read the maximum from TEST_COST, never hard-code it

The judge passes the test's own point value to the checker in the environment, along withEOLYMP=1, INPUT_FILE, OUTPUT_FILE, ANSWER_FILE, TEST_ID, TEST_INDEX andTEST_GROUP:

static double test_cost() {
    const char *cost = getenv("TEST_COST");
    if (!cost || !*cost)
        quitf(_fail, "TEST_COST is not set, so the subtask's point value is unknown");
    return atof(cost);
}

Reading the maximum from TEST_COST keeps subtask weights in the testset configuration,where they belong, instead of baked into checker source.

7.2 Eolymp reads ABSOLUTE points, and clamps to the test's cost

Eolymp scans the checker's stdout for points <float> and treats it as absolute points,then clamps to the test's own cost.

A checker returning a percentage therefore fails silently: it pays full marks on everytest whose cost is below the percentage, and nothing on the rest. Measured on IOI 2010 Maze:points 90.9090909091 on a test worth 11 was awarded 11, and the problem scored 11 of110 — which reads like nine dead subtasks, not a units error.

A checker is not proved correct by one subtask paying out. Verify on a test whose cost isboth above and below the value you return.

7.3 A CMS fork of testlib prints something else

Stock testlib's quitp(x, msg) prints points x. A CMS fork (e.g. APIO'sbike/checker/testlib.h) prints a bare number, treats the argument as a fraction, andclamps small values up to 1e-5 to avoid zeros in CMS. Shipping the fork compiles fine andthen scores nothing Eolymp can read.

Port the helpers onto stock testlib instead — registerChecker, readSecret,readGraderResult are about fifteen lines between them — and multiply the fraction by thetest cost. (readSecret/readGraderResult show up in IOI-style grader problems.)


8. Floating-point answers

Rare, because olympiad answers are overwhelmingly integral.

double j = ans.readDouble(), p = ouf.readDouble(-1e9, 1e9, "answer");
if (!doubleCompare(j, p, 1e-6))
    quitf(_wa, "expected %.10f, found %.10f, error = %.10f", j, p, doubleDelta(j, p));

doubleCompare(expected, found, eps) handles absolute or relative error, which is whatstatements mean. Do not hand-roll fabs(a-b) < eps — it fails at large magnitudes. UsereadStrictDouble(min, max, minDigits, maxDigits) when the statement fixes the outputformat. Never let an unbounded readDouble() take contestant input: nan and inf parse.


9. Messages

The message is the only artifact a human sees when a verdict is disputed, or when the model solution starts failing — and onCodeforces contestants see it during hacking.

quitf(_wa, "wrong at position %d: expected %d, found %d", i, want, got);   // good
quitf(_wa, "wrong answer");                                               // useless

Name the position, the expected value and the found value. Never print to stdout yourself —the verdict machinery owns it.


10. Review checklist

BLOCKER breaks judging · MAJOR hides defects · MINOR style.

Structure

  • [ ] BLOCKER registerTestlibCmd(argc, argv) present

  • [ ] BLOCKER checker type is PROGRAM, not LEGACY_PROGRAM (argument order differs)

  • [ ] MAJOR a built-in type (TOKENS/LINES/QUERY_RESULTS) would not have sufficed

  • [ ] BLOCKER every path ends in a quit* — no falling off the end of main

  • [ ] BLOCKER composite answer ⇒ readAns(InStream&) used for both streams

  • [ ] BLOCKER readAns(ans) called before readAns(ouf)

  • [ ] BLOCKER stream.quitf inside readAns, never plain quitf

  • [ ] MAJOR trailing-content policy deliberate (ouf.seekEof() or explicitly lenient)

Trusting the contestant

  • [ ] BLOCKER every ouf.read* producing a count/index/length is bounded

  • [ ] BLOCKER no container sized, and no array indexed, from an unchecked value

  • [ ] MAJOR every read* names its variable

  • [ ] MAJOR no unbounded loop over untrusted output

Verdict semantics

  • [ ] BLOCKER _fail only for jury-side impossibilities, never contestant errors

  • [ ] BLOCKER contestant beating the jury ⇒ _fail, not _ok

  • [ ] MAJOR feasibility and optimality both checked, for optimisation problems

  • [ ] MAJOR messages name position, expected and found

Scoring (if partial)

  • [ ] BLOCKER returns absolute points, not a percentage or fraction

  • [ ] BLOCKER the maximum comes from getenv("TEST_COST"), not a hard-coded constant

  • [ ] BLOCKER stock testlib, not a CMS fork

  • [ ] BLOCKER quitp(...), not _pc(...)

  • [ ] MAJOR verified on a test whose cost is both above and below the returned value

  • [ ] MAJOR full marks, partial and zero all reachable