Interactors with testlib.h

When this page applies. This is the testlib version of Write an interactor, which uses eolymp.h and is the default for a new interactor. Use this page only when:

  • the interactor being changed already uses testlib — it includes testlib.h or calls registerInteraction (read it first with get_interactor). Change it in testlib; do not port it to eolymp.h unless the requester asks for exactly that; or

  • the requester explicitly asks for a testlib interactor.

The rule, and why, is in eolymp.h reference.

1. What an interactor is

A program that talks to the contestant's process over a pipe, in real time. It reads thetest from inf, answers the contestant's queries, and decides whether the session wascorrect.

  inf  ─────►  interactor  ◄──── stdin ────  contestant
                    │      ────► stdout ───►
                    └──► tout ──►  (read later by the checker)

The mapping is not the same as a checker's, and the second argument is the trap:

stream

comes from

inf

argv[1] — the test input

tout

argv[2] — the interactor's own log, which the checker later reads as its ouf

ouf

stdin — the live pipe from the running solution

ans

argv[3] — the jury answer, when the problem has one

argv[2] is not the contestant's output. The contestant's output arrives on stdin as ouf;argv[2] is where the interactor writes what it wants the checker to see.

ans is available and is worth using when the verdict depends on comparing against a correctsolution's product. For a guessing game it is not needed — the secret is in inf.

The problem's type must be INTERACTIVE, and the interactor is configured withupdate_interactor.

Its exit-code contract differs from a checker's: 0 means the interaction succeeded andthe checker then grades the test; 1 fails that test as a wrong answer; any other codefails the whole submission as a system error.

Three things make it different from a checker, and all three are ways to get it wrong:

  1. It runs concurrently with the contestant. A missing flush is a deadlock, not a wronganswer.

  2. It is the contestant's only source of truth. Every reply it sends is part of theproblem statement's contract.

  3. It usually does not score. Scoring belongs to a paired checker, which reads what theinteractor writes to tout (§6).


2. The skeleton

#include "testlib.h"
using namespace std;

int main(int argc, char* argv[]) {
    registerInteraction(argc, argv);          // always first

    int n = inf.readInt();                    // the hidden test
    int secret = inf.readInt();

    cout << n << endl;                        // endl FLUSHES — that is the point

    int queries = 0;
    const int LIMIT = 20;

    while (true) {
        string cmd = ouf.readToken("!|\\?", "command");

        if (cmd == "!") {                     // final answer
            int guess = ouf.readInt(1, n, "answer");
            if (guess != secret)
                quitf(_wa, "wrong answer: said %d, actual %d", guess, secret);
            tout << queries << endl;          // hand the query count to the checker
            quitf(_ok, "guessed in %d queries", queries);
        }

        int q = ouf.readInt(1, n, "query");
        if (++queries > LIMIT)
            quitf(_wa, "exceeded %d queries", LIMIT);

        cout << (q < secret ? "GREATER" : q > secret ? "LESS" : "EQUAL") << endl;
    }
}

Runtime: store it as cpp:20-gnu14. Runtime ids are language:version-toolchain;a bare cpp17 or cpp is rejected — see testlib reference.


3. Flushing: the deadlock rule

Every write to the contestant must be flushed. This is not a performance question — an unflushed reply means the contestant blocks waiting for input the OS is stillbuffering, while the interactor blocks waiting for a query that will never come. Bothprocesses hang until the judge's time limit kills them, and the verdict is a meaningless TLEagainst a correct solution.

cout << x << endl;      // endl flushes
fflush(stdout);         // after printf
cout << flush;

endl is correct here — the flush is the point. This is the exact opposite of the rulefor generators, where endl is a performance bug. Same token, opposite advice, because theconstraint is different.

Never use printf without fflush(stdout). Never mix printf and cout unlesssync_with_stdio is left on. And never set ios_base::sync_with_stdio(false) in aninteractor without auditing every write path.


4. Never trust the contestant's stream

ouf is the contestant's live output — arbitrary bytes from a process that may be buggy,adversarial, or already dead.

int k = ouf.readInt();                 // BAD — accepts 2^31-1, then you loop k times
int k = ouf.readInt(1, n, "k");        // GOOD

Specific hazards, in order of how often they bite:

  • Unbounded counts. A query of the form "I will now send k values" with k unboundedlets a bad submission make the interactor allocate or loop unboundedly.

  • Do not finish with ouf.readEof(). It asserts EOF is the very next character andfails with "Expected EOF" when the solution left a trailing newline. Useif (!ouf.seekEof()) quitf(_wa, "extra output after the final answer"); — seekEof skipswhitespace first.

  • The contestant exits early. Reading from a closed pipe. Testlib turns this into averdict rather than a crash, but if you loop while (true) with no seekEof guard and noquery limit, behaviour depends on testlib's EOF handling rather than on your intent.

  • A failed read is already a verdict. If any ouf.read* fails in an interactor — wrongtype, out of bounds, unexpected EOF — testlib ends the run as Wrong Answer on its own.So a bounded read is not only a guard against your own crash, it is the verdict.

  • Malformed tokens. Use ouf.readToken("!|\\?", "command") with an explicit patternrather than reading a string and comparing — it produces a proper verdict on garbageinstead of falling through your if/else chain into undefined behaviour.


5. Query limits

Every interactive problem with a query budget must enforce it in the interactor — thestatement saying "at most 20 queries" is not enforcement. This is the single most oftenomitted check, especially on tasks ported from CMS function-call graders, where the graderenforced the budget in-process and the enforcement has to be rewritten.

if (++queries > LIMIT)
    quitf(_wa, "exceeded the query limit of %d", LIMIT);

Two details worth getting right:

  • Count the query when it is read, not when it is answered. Otherwise the lastover-budget query still gets a truthful reply, which can leak the answer.

  • Decide whether exceeding the limit is _wa or a scored zero. For a subtask-scoredproblem it is usually a zero for that subtask via tout, not a hard _wa — see §6.

If the limit is adaptive (the problem promises the interactor is adaptive), say soexplicitly in a comment and make sure the interactor is genuinely adaptive: it must beconsistent with every answer still compatible with the queries so far, not secretlycommitted to one hidden value. An "adaptive" interactor that actually fixes the answer upfront is a statement bug.


6. tout and the interactor/checker split

The interactor writes to tout; the paired checker scores. That is the shape of everyscored interactive problem:

  • the interactor validates the protocol and writes a summary to tout (query count,score fraction, per-subtask flags)

  • the checker reads tout as its ouf and produces the verdict and points

// interactor
tout << queries << endl;
quitf(_ok, "session valid");

// checker  (ouf is the interactor's tout)
int queries = ouf.readInt(0, 1000000, "queries");
if (queries <= 10) quitp(100, "%d queries", queries);
else if (queries <= 20) quitp(50, "%d queries", queries);
else quitf(_wa, "too many queries: %d", queries);

The checker must consume everything the interactor wrote. tout reaches the checker asits ouf, and an unread remainder is reported as "Extra information in the output file" —a presentation error on a session that was actually fine. Either read every value theinteractor writes, or drain the rest:

while (!ouf.seekEof()) ouf.readToken();

That is why a scored interactive problem needs a checker even when the interactor alreadydecided everything: without one, nothing consumes tout.

Why the split exists: the interactor exits as soon as the session ends, but scoring maydepend on the whole session. Keeping the score in the checker also means the scoring rulecan be fixed without touching the protocol.

The platform takes a run's score from the checker and has no scoring channel for theinteractor at all. So:

  • A pass/fail interactive problem may run with no checker. The interactor's own exit codedecides the test, and disabling the checker is a supported configuration.

  • A partially scored one needs a checker, because points can only come from there — andthe checker is then also what consumes tout.

Everything about checker scoring applies to that checker: it returns absolute points,read from TEST_COST, clamped to the test's own cost. A percentage returned here failsexactly as silently as it does anywhere else.


7. The throughput ceiling — a real porting hazard

Eolymp runs the interactor as a separate process, so every query is a pipe round trip.Measured locally at roughly 150,000 queries/second, and a fast Linux judge is the sameorder.

Olympiads that link their grader into the contestant's binary have no such cost and sizetheir tests accordingly. CEOI 2026 "Treasure Hunt" ships 100,000 hunts per test — about22 million queries in the largest file — against a stated 8 s limit. Measured on Eolymp:a mid-size file took 28 s and the largest was still running after ten minutes.

Before porting a function-call interactive problem to a process-based interactor,compute:

units × queries-per-unit ÷ 150,000  >  time limit ?

If it exceeds the limit, the only lever that does not change the problem is fewer units pertest. Keep every test file, subtask and parameter; a scoring rule that is a minimum overunits is the same quantity measured on a smaller sample. Take an even spread rather than aprefix, so a file that front-loads corner cases keeps them. The cost is sensitivity: asolution failing on a small fraction of units is less likely to be caught.


8. Determinism and the hidden state

The interactor reads its secret from inf, which the generator produced and the validatorvetted. It must not invent state:

  • No rand(), no rnd, no clock. An interactor that randomises makes the verdictirreproducible, so a failing submission cannot be debugged and a rejudge can changeresults.

  • Everything the contestant can learn must come from inf.

  • If the problem is adaptive, the adaptation must be a deterministic function of the queryhistory and inf.


9. Review checklist

BLOCKER breaks judging · MAJOR hides defects · MINOR style.

Protocol

  • [ ] BLOCKER registerInteraction(argc, argv) present

  • [ ] BLOCKER every write to the contestant is flushed (endl, flush, or fflush)

  • [ ] BLOCKER no sync_with_stdio(false) without auditing every write

  • [ ] BLOCKER query limit enforced in code, counted on read

  • [ ] MAJOR EOF / early contestant exit handled deliberately

  • [ ] MAJOR finishes with !ouf.seekEof(), not ouf.readEof()

  • [ ] MAJOR malformed commands produce a verdict, not fall-through

Trusting the contestant

  • [ ] BLOCKER every ouf.read* producing a count/index is bounded

  • [ ] BLOCKER no container sized or array indexed from an unchecked value

  • [ ] MAJOR the reply cannot leak information beyond the stated protocol

  • [ ] MAJOR an over-budget query is not answered truthfully before the verdict

Scoring

  • [ ] MAJOR tout written on every terminating path that should score

  • [ ] BLOCKER the paired checker consumes all of tout, or drains the remainder

  • [ ] MAJOR the paired checker reads tout and returns absolute points

  • [ ] MAJOR full marks, partial and zero all reachable and verified

Determinism

  • [ ] BLOCKER no randomness, no clock, no state outside inf and the query history

  • [ ] MAJOR an "adaptive" interactor is genuinely adaptive, not secretly fixed

Feasibility on this platform

  • [ ] MAJOR total queries ÷ 150,000 is comfortably under the time limit (§7)


10. How to test an interactor

An interactor cannot be exercised by feeding it a file — it needs a live partner, and theplatform wires the pipes. So the only real tests are submissions:

  1. Submit the model solution. It must score full marks. This is the only end-to-end test.

  2. Submit deliberately broken solutions and confirm each verdict:

    • one that exceeds the query limit → the limit verdict, not TLE

    • one that answers wrongly → _wa

    • one that prints garbage → a verdict, not a hang

    • one that exits immediately without querying → a verdict, not a hang

    • one that never flushes → should TLE, and this confirms the contestant side of theflush contract is what breaks, not your interactor

  3. Time the largest test with the model solution and compare against §7.

The hang cases are the ones worth the effort: a checker bug produces a wrong verdict, but aninteractor bug produces a hung judge, and hangs are found only by trying them.