{
  "markdown": "# @exercode/problem-utils\n\n[![Test](https://github.com/WillBooster/exercode-problem-utils/actions/workflows/test.yml/badge.svg)](https://github.com/WillBooster/exercode-problem-utils/actions/workflows/test.yml)\n[![semantic-release](https://img.shields.io/badge/%20%20%F0%9F%93%A6%F0%9F%9A%80-semantic--release-e10079.svg)](https://github.com/semantic-release/semantic-release)\n[![wbfy](https://img.shields.io/badge/wbfy-20.23.1-1e90ff.svg)](https://github.com/WillBooster/shared/tree/main/packages/wbfy)\n\n:100: A set of utilities for judging programs on Exercode (https://exercode.willbooster.com/).\n\n## Packages\n\n| Package                           | Purpose                                                                                                  |\n| --------------------------------- | -------------------------------------------------------------------------------------------------------- |\n| `@exercode/problem-utils`         | CLI, validation, result types, and stdio/command/GUI/evaluation presets; no browser or AI SDK dependency |\n| `@exercode/problem-utils-browser` | Puppeteer Chromium lifecycle, browser judging, and screenshots                                           |\n| `@exercode/problem-utils-llm`     | LLM judging and AI SDK providers                                                                         |\n\nInstall only the packages the problem uses. Import `llmJudgePreset` from\n`@exercode/problem-utils-llm`. Browser judges receive Puppeteer's native `Page`:\n\n```ts\nimport { DecisionCode } from '@exercode/problem-utils';\nimport { browserJudgePreset } from '@exercode/problem-utils-browser';\n\nawait browserJudgePreset({\n  testCases: [\n    [\n      'heading',\n      async (page) => ({\n        decisionCode:\n          (await page.$$eval('h1', (elements) => elements.length)) === 1\n            ? DecisionCode.ACCEPTED\n            : DecisionCode.WRONG_ANSWER,\n      }),\n    ],\n  ],\n});\n```\n\nThe browser preset serves the submitted directory, runs checks in order on one page,\nprints each result, stops on the first non-accepted result, and closes the browser even\nwhen a check throws. Checks return learner-facing verdicts; uncaught harness errors\npropagate to the caller. Use `timeoutMs`, `viewport`, `contextOptions`, and `launchOptions` to set\nproblem-specific requirements. `screenshotOnFailure` attaches a full-page image to a\nnon-accepted result; capture failures are recorded in `stderr` without replacing the verdict.\n`launchBrowser`, `createBrowserPage`, and `captureScreenshot` are also exported for harnesses\nthat manage their own HTTP server or test loop. Use `await createBrowserPage(browser)`\n(or pass a browser context) to create a native Puppeteer page with the same dialog handling\nas the presets. These pages dismiss dialogs by default. Registering a `dialog` listener takes\nover handling: the listener must call `dialog.accept()` or `dialog.dismiss()`, even when it\nonly inspects the message. Asynchronous handlers retain control until they handle the dialog.\n\nThe browser package depends on `puppeteer`. Before running browser judges or PDF export,\ninstall its version-matched Chrome headless shell explicitly when package installation scripts are disabled:\n\n```sh\nbun run exercode-browser browsers install chrome-headless-shell\n```\n\n`exercode-browser` runs this package's Puppeteer CLI, independently of an application's\nE2E test version. Docker/CI should install `unzip` before this command and Chrome's OS\nlibraries at image build/setup time. Install the fonts required by the course content in that environment.\n\n## HTML comparison\n\n`htmlJudgePreset({ solutionDirectoryPath, requiredFiles?, compareDom?, textNormalizationPattern? })` compares the submitted\npage with the specified model-answer directory. It reports `snapshot_body` first,\nthen `screenshot`, stopping on the first difference. Set `compareDom: false` for exercises that grade only the rendered appearance. The DOM comparison ignores\ncomments and normalizes text whitespace and attribute order. `textNormalizationPattern` overrides\nthe regular-expression source used to replace text-node matches with spaces before trimming;\nit defaults to `\"\\\\s+\"`. Set it when a course requires a different text comparison rule. Screenshots render\nformatted HTML at 800×600, with CSS animations disabled and fonts loaded; a difference\nincludes both PNG files. HTML decoding honors a BOM or declared HTTP/meta charset, defaulting to UTF-8 when\nnone is declared; formatted responses explicitly use UTF-8. If either document cannot\nbe decoded or formatted, both are rendered raw.\nEach check uses fresh, separate browser contexts for both answers. Pixel comparison requires deterministic\npage content; JavaScript timers, random content, and animated images are not frozen. Missing required files are reported before starting Chromium.\n\nFor isolated judging, keep shared assets inside the problem directory.\nBoth directories can use the nearest ancestor's `assets` directory, through `assets/`\nor directly from the served root. Local assets are merged with that shared directory;\nsubmission files and links take precedence, followed by local assets. A directory\nsymlink overrides that entire directory rather than merging shared files into its target. Source files and\nlinked directories are left unchanged. The temporary\nserved directories and browser are closed after judging. The directory helpers leave\nprocess signals to the host. Hosts must provide and remove a per-run `TMPDIR` when a\nharness terminates before disposal; isolated CLI checks do this automatically for\nSIGINT, SIGTERM, and SIGKILL. Long-lived hosts can await disposal in their own\nshutdown handlers. `captureHtmlBodySnapshot`,\n`captureHtmlScreenshotPair`, and `createHtmlServedDirectory` expose the same operations\nfor custom checks. The screenshot pair takes two `{ page, url }` targets and returns\nPNGs in that order, formatting both documents or neither. Give those pages matching\nviewport options and separate fresh browser contexts. See the [HTML example](example/web_page_comparison/judge.ts).\n\n## Custom browser pages and lifecycle hooks\n\n`browserJudgePreset` accepts `directoryPath` to serve an assembled exercise directory,\n`entryPath` to select its initial page, and `navigationOptions` for native Puppeteer\nnavigation settings. `initializePage` runs before navigation, so it can register console\nand page-error listeners. `afterTests` runs after the checks, including a failing verdict,\nwhile the page remains open. An exception in initialization, navigation, or a check\npropagates to the caller and skips `afterTests`; browser cleanup still runs.\n\n## Spring Boot courses\n\nImport `springBootJudgePreset` from `@exercode/problem-utils/presets/springBoot` and pass\n`problemDirectoryPath` plus an async `evaluate` callback. The preset builds the problem's\nMaven project offline, starts its Spring Boot JAR on port 59000, and invokes the problem's\n`judge.ts --evaluate`. It preserves the course build, startup, and evaluation limits\n(90, 60, and 60 seconds) and emitted screenshot metadata. The host provides Java, Maven,\nBun, cached dependencies, an exclusive execution slot, and interrupted-run cleanup.\nEvaluators can build their request URLs with `buildSpringBootUrl` from the same core subpath.\nBrowser evaluation can use `launchBrowser` and `captureTomcatScreenshots` from the browser\npackage; the Spring Boot preset itself adds no browser dependency to core.\n\n## JavaScript and Tomcat courses\n\n`javascriptJudgePreset(problemDirectoryPath, options)` from `@exercode/problem-utils-browser`\nruns `main.mjs` or `main.js` in Chromium and compares console output with `test_cases/*.out`.\nThe `.in` files contain browser setup JavaScript. Set `initializeAndVerifyDom: true` to\ncall the setup's `initializeTest` and `verifyDom` hooks, and `waitForConsoleIdle: true`\nfor exercises whose asynchronous console output must settle before comparison.\n\n`javascriptDomJudgePreset(problemDirectoryPath)` supports DOM exercises whose `.in`\nfiles define `initializeTest`, `verifyDom`, or `window.test`. Setup scripts retain global\ndeclarations for submitted code and verification hooks. The preset exposes top-level\nfunction declarations to those callbacks, captures console output, and stops at the\nfirst failing case. Both JavaScript presets require a host-enforced overall timeout:\nthe console-idle heuristics do not bound programs that keep emitting output.\n\n`tomcatJudgePreset` and `buildTomcatUrl` are available from\n`@exercode/problem-utils/presets/tomcat`, without browser dependencies. The preset\naccepts `problemDirectoryPath`, `jspDirectory` (`''` or `'WEB-INF/jsp'`), optional\n`forbiddenTexts`, and an asynchronous `evaluate` callback. It builds the problem's\n`pom.xml` with Maven offline, serves the `judge` application on port 59000, and invokes\nthe same `judge.ts` with `--evaluate`. The host must provide Maven, GNU time, GNU\ntimeout, Bun, `CATALINA_HOME`, and exclusive access to that port. Build and evaluation\nlimits are 60 and 30 seconds respectively.\n\nBrowser evaluations can use native Puppeteer pages together with\n`captureTomcatScreenshots` or `verifyTomcatHtml` from the browser package. The latter\ncompares document markup with whitespace removed and records screenshots for the\nverdict. `verifyTomcatPath(page, endpoint)` waits up to five seconds for the endpoint's\npath and parsed document, reporting the expected and current paths in Japanese if navigation fails;\nquery strings do not affect the check. All interrupted-run cleanup remains the host's\nresponsibility. `requirePageElement(page, selector)` immediately returns a native\nPuppeteer element handle or reports the missing selector in Japanese; callers use\nthe handle's native actions for course-specific interactions.\n\n## GUI programs\n\n`guiCommandJudgePreset` from `@exercode/problem-utils/presets/guiCommand` runs a program on Xvfb,\ncaptures every top-level window, waits `screenshotWaitSeconds` (0.3 by default) before the next\ncapture (so captures are that wait plus the time a capture takes apart), and stops the\nprogram once `stopDetectionThreshold` (5 by default) consecutive captures are identical. The\nproblem's `test` receives the last capture of each window as PNG files in `runResult.screenshots`.\nThe host must provide `Xvfb`, `maim`, `xdotool`, and `xwininfo`.\n\nSet `recordsAnimation: true` for programs a grader has to watch moving, such as animations:\n\n```ts\nimport { DecisionCode } from '@exercode/problem-utils';\nimport { guiCommandJudgePreset } from '@exercode/problem-utils/presets/guiCommand';\n\nawait guiCommandJudgePreset(import.meta.dirname, {\n  recordsAnimation: true,\n  test: ({ runResult }) => ({\n    // A program that is still animating at the time limit is fine; one that shows nothing is not.\n    decisionCode:\n      runResult.stopReason === 'timeout' && runResult.screenshots.length === 0\n        ? DecisionCode.TIME_LIMIT_EXCEEDED\n        : DecisionCode.ACCEPTED,\n    outputFiles: [...runResult.screenshots, ...runResult.recordings],\n  }),\n});\n```\n\n- `runResult.recordings` holds an animated PNG (`<window name>_<window id>_recording.png`) of each\n  window that kept changing during the run, assembled from the captures above and looping forever.\n  Browsers play it wherever they show a PNG, so Exercode displays it like a screenshot. A window that\n  only appeared and was painted has no recording, only its screenshot.\n- A run that is still changing at the time limit reaches `test` with `stopReason: 'timeout'` instead of\n  being reported as `TIME_LIMIT_EXCEEDED`, so `test` decides the verdict (e.g. a time limit exceeded\n  when nothing was captured). Without `recordsAnimation`, such a run never reaches `test`.\n- The recordings of one run take at most `MAX_GUI_RECORDING_BYTES` (2 MiB) in total: the largest one\n  loses every other frame until they fit, and one that does not fit even with two frames is left out, so\n  decide a timed-out run by `runResult.screenshots` rather than by the presence of a recording. A lower\n  `screenshotWaitSeconds` records more smoothly; raise `stopDetectionThreshold` with it, and for programs\n  that pause longer than the two multiplied.\n\n## PDF export\n\n```typescript\nimport { markdownToPdf } from '@exercode/problem-utils-browser/pdf';\n\nconst pdf = await markdownToPdf(markdown, {\n  assetDirectoryPath: import.meta.dirname,\n  pdfOptions: { format: 'A4' },\n});\nawait Bun.write('material.pdf', pdf);\n```\n\nPDF assets are served on IPv4 loopback, and encoded paths cannot escape the asset directory. Intentional asset symlinks remain usable. Core also exports `startLocalHttpServer` for callers that need the same local-only server; await its startup before using its address.\n\nThe PDF entry point removes YAML mapping frontmatter, renders Markdown with syntax highlighting and CJK-friendly emphasis, and resolves relative\nimages against `assetDirectoryPath`, and waits for fonts and images before printing. Missing or invalid images leave\nbrowser placeholders without preventing the document from exporting.\nPass `mermaidScriptPath` pointing to a Mermaid browser bundle to render diagrams,\n`css` to customize styling, and `pdfOptions` for native Puppeteer PDF settings.\nThe defaults use screen media, A4 paper when no custom dimensions are supplied, printed backgrounds, and margins of\n30 mm top/bottom, 40 mm right, and 20 mm left. PDF rendering dependencies are loaded\nthrough this subpath; importing the browser judging entry point does not load them.\n\n## CLI\n\nThe package ships an `exercode-problem` command for problem authors (run it with `bun x` in a repository that depends on `@exercode/problem-utils`):\n\n```bash\n# Judge all model answers of all problems (directories containing problem.md or <id>.problem.md) under a directory.\n# model_answers/* must be fully accepted; model_answers.fails/* must fail at least one test case.\nbun x exercode-problem                # everything under the current directory\nbun x exercode-problem courses/foo    # everything under courses/foo\nbun x exercode-problem --only a_plus --skip gui_ --concurrency 2\n\n# Judge one answer directory of the problem in the current directory.\nbun x exercode-problem judge model_answers/python\nbun x exercode-problem judge model_answers/python '{ \"language\": \"python\" }'\n\n# Debug one answer directory of the problem in the current directory.\nbun x exercode-problem debug model_answers/python '{ \"stdin\": \"1 2\" }'\n```\n\n`judge` and `debug` run a custom `judge.ts` / `debug.ts` when the problem has one, and apply `stdioJudgePreset` / `stdioDebugPreset` otherwise, mirroring the Exercode server. The debug fallback applies only to standard problems: a problem with a custom `judge.ts` needs its own `debug.ts`, and `exercode-problem debug` fails with a message otherwise (the server likewise reports debug as unsupported there).\n\nThe all-problem check judges serially by default because time limits are measured in wall-clock time; pass `--concurrency <n>` to parallelize when the checked problems are not timing-sensitive.\n\nA standard stdin/stdout problem must NOT commit a `judge.ts` or `debug.ts` that is identical to the default stdio harness: the absence of `judge.ts` marks the problem as standard, and committed copies would drift from the server's defaults. The CLI rejects such files; a file kept intentionally (e.g. to demonstrate the default harness) can add an explanatory comment to be treated as custom.\n\n## Validators\n\nThe CLI also validates learning-material files without running any program, mirroring the checks the Exercode importer applies:\n\n```bash\n# Validate problem directories (problem.md frontmatter, test_cases, model_answers, templates, judge.ts / debug.ts).\nbun x exercode-problem validate-problem <problemDir>...\n# Validate a course directory (course.yaml, lecture materials with embedded questions, problem references).\nbun x exercode-problem validate-course <courseDir> [--problems-dir <dir>]\n# Validate a contest (*.contest.yaml) file.\nbun x exercode-problem validate-contest <contestYamlPath> [--problems-dir <dir>]\n```\n\nEach target prints `OK` or `NG` followed by its errors and warnings; the command exits 1 when any target has an error. `--problems-dir` points to the directory holding the referenced problems; a course is always searched at any depth, since Exercode links a material only to the problems inside its course, so for a course the option must name a directory inside it. The validators are also exported (`validateProblemDirectory`, `validateCourseDirectory`, `validateContestFile`, `validateMaterialFile`).\n\n## Agent skills\n\nThe [`skills/`](skills/) directory holds skills for AI coding agents that author and review Exercode learning content: `generate-learning-content` (entry point), `generate-course-materials`, `generate-judge-problems`, `generate-judge-contest`, `review-learning-content`, and `setup-exercode-course-repository`. Agents working in this repository load them through the symlinks under `.claude/skills/`; install them elsewhere with the [skills](https://github.com/vercel-labs/skills) CLI:\n\n```bash\nbun x skills add WillBooster/exercode-problem-utils --agent claude-code --agent codex\n```\n\n## Test cases\n\nA problem keeps its test cases under `test_cases/`. A test case id is the shared name of the following entries, and each entry is optional:\n\n| Entry          | Meaning                                                                                                                                           |\n| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |\n| `<id>.in`      | Standard input. Omit it (or leave it empty) when the program reads nothing.                                                                       |\n| `<id>.out`     | Expected standard output.                                                                                                                         |\n| `<id>.fin/`    | Files copied into the working directory before the run (input files).                                                                             |\n| `<id>.fout/`   | Expected output files, compared with the files of the same relative paths in the working directory.                                               |\n| `_shared.fin/` | Files copied into the working directory before every test case.                                                                                   |\n| `<id>.json`    | Configuration for a custom `judge.ts` that reads it itself; the presets ignore it (the Judge server lists it as a test case of the custom judge). |\n\nA test case whose id contains `example` (the judge server's rule, e.g. `example_1` or `01_example_small`) is an example shown to learners; every other case is hidden.\n\nStandard output and text files are compared as space-separated tokens: consecutive white spaces count as one separator, and an expected token that contains a decimal point and parses as a finite number (e.g. `3.14`, but not `1`, `1e-3` or `1.0e309`) accepts a value within an absolute or relative error of `1e-6`. A file is text when it is valid UTF-8 without NUL bytes; other files (e.g. images) must match byte for byte. A received file larger than 8 MiB counts as not produced, an expected file larger than 8 MiB is an authoring error (the case is reported as a runtime error), and a file larger than 1 MiB is left out of the reported pair. When a file differs, the result carries `<name>_expected.<ext>` and `<name>_received.<ext>` so Exercode can show both (Exercode decides per test case whether a learner may see them, as it does for expected stdout).\n\nHow a missing expectation is treated depends on the harness:\n\n- `stdioJudgePreset` (the default for problems without `judge.ts`) requires `<id>.out` or a non-empty `<id>.fout/` for every test case, so a standard problem cannot accept a run without checking it. A problem whose `problem.md` declares `requiredOutputFilePaths` or `isManualScoringRequired` is exempt, because every test case is judged by those instead; code rules (`requiredRegExpsInCode` etc.) and `requiredSubmissionFilePaths` check the submission once, in addition to the output comparison, and do not exempt.\n- `commandJudgePreset` without a `test` option checks whatever expectations exist, and a test case with neither only has to run within the limits. A `test` option replaces that comparison; it receives `testCase.output`, `testCase.fileOutputPath` and `cwd` and can call the exported `judgeAgainstExpectations` (or `compareStdoutAsSpaceSeparatedTokens` and `compareExpectedOutputFiles`). A custom `readTestCases` may return any test case type with `id` (plus optional `input` and `fileInputPath`); the default verdict judges any case that exposes a string `output` or `fileOutputPath`, which the default reader's `CommandTestCase` does.\n- `guiCommandJudgePreset` passes the expectations to the problem's `test`, which decides everything.\n- `llmJudgePreset` runs no program, so it copies no `.fin/`; it hands `<id>.in` as the prompt input and the whole entry (including `fileOutputPath`) to the problem's `test`.\n- `stdioDebugPreset` by default copies the answer directory to a temporary directory and builds and runs it there together with `_shared.fin/` and the first example case's `.fin/` (the case's files win; hidden cases' inputs are never handed to learner code), so the answer directory is left untouched (its `node_modules`, and those of its ancestors, are linked into the copy rather than duplicated); files the program writes are reported only through `requiredOutputFilePaths`. Callers that already own a disposable answer directory can pass `{ disposableWorkingDirectory: true }` as the second argument to build and run directly there; they must remove it themselves, including when the harness is terminated.\n- `evaluationJudgePreset` does not use `test_cases/`.\n\n`readTestCases` is exported for harnesses that enumerate `test_cases/` themselves.\n\n## Measurements\n\n`stdioJudgePreset`, `stdioDebugPreset`, and `commandJudgePreset` (with its default runner; a custom `runCommand` decides whether its results carry `cpuTimeSeconds`) run a program under GNU time (`/usr/bin/time` on Linux, `gtime` on macOS) and report its wall time (`timeSeconds`), user plus system CPU time (`cpuTimeSeconds`), and peak resident set size (`memoryBytes`, at least the footprint of the GNU `timeout` wrapper the program runs under, about 1 MiB) in every test case result; `guiCommandJudgePreset` reports the wall time and the peak resident set size, and `llmJudgePreset` the wall time only. Time limits are judged by wall time. The CPU time is recorded even for a run that exceeded its limit, so a judge server sharing CPUs between programs can tell a program that used up its limit (its CPU time, summed over its threads, reaches the limit, so it would exceed it on any single CPU) from one that may only have waited for a CPU (its CPU time stays below the limit) and re-run just the latter alone.\n\n`browserJudgePreset` forwards measurements returned by each check; it does not measure elapsed time, CPU time, or memory automatically. Its `timeoutMs` sets the page default timeout for navigation, waiting and shared screenshot capture. Other native operations follow Puppeteer’s own timeout behavior; this option does not bound the complete check. Checks that require their own measurements or time-limit verdicts must provide them explicitly. The hosting judge can independently limit the overall harness process.\n",
  "bytes": 23560,
  "sha": "edf997b168e45cd4ca097175aaf4856cabdc0130ac24fe18b4b7e54f5c8a693c",
  "repo_slug": "willbooster/exercode-problem-utils",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_willbooster_exercode_problem_utils_revie_0d10e4de/readme"
}