{
  "markdown": "# AIQ\n\nAIQ records fixed-fixture AI and agent benchmark results. The repository\ncontains a Rust runner, a Rust verifier, a Next.js application, the public AIQ\nCore catalog, and one declarative PostgreSQL schema.\n\nAIQ production is live at [aiq.wiki](https://aiq.wiki). The personal Vercel\nscope `acgbox` hosts project `aiq`. The personal Supabase organization `ACG Box`\nhosts project `aiq` on PostgreSQL 17.6 with reference\n`xxnszykaeapolqdnhalx`. The personal Cloudflare account that owns the\n`aiq.wiki` zone owns DNS handoff. Production uses the private Storage buckets\n`aiq-submission-packages` and `aiq-runner-artifacts`.\n\nThe only supported production tuple is AIQ Core `1.1.0`, task scorer `1.0.6`, aggregate\nscoring `1.0.8`, and measurement `2.0.0`. Do not publish, preserve online,\nmigrate, or display a legacy tuple as production evidence. Production must\nremain without an Official AIQ 2.0 publication until a fresh complete,\nnon-synthetic, signed 17-by-72 calibration is replayed under policy v2 to\nestablish the fixed item bank and admission v3, and\na separate fresh 17-by-72 Official package passes native verifier replay and\nall release gates.\n\nFormal calibration and Official model and evaluator work has no\nbenchmark-enforced wall-time, step, tool-call, aggregate-evaluator, or\nper-check deadline. The runner separately measures model and evaluator elapsed\ntime, agent steps, tool calls by type, tokens, and estimated cost. These values are auxiliary\nevidence only and cannot change task score, AIQ, quality, strict pass, interval,\neligibility, or ranking. Functional preflight and hard safety boundaries remain\nseparate. A safety, runtime, provider, or infrastructure termination produces a\nnull semantic score, never a semantic zero.\n\nAll earlier bounded or deadline-bearing runs remain immutable failed release\nevidence. They cannot be relabeled, composed with selected reruns, or published\nunder the active tuple.\n\n## Product contract\n\n- Repository source targets AIQ Core `1.1.0`, with 72 private controlled tasks\n  in ten domains. Task evaluation stays at `1.0.6`; aggregate scoring is\n  `1.0.8`.\n- Every formal task encodes `wall_seconds: null`, `max_steps: null`, and\n  `max_tool_calls: null`. Controlled evaluator configuration uses\n  `aiq.evaluator-config.v2` with `completion_policy: natural_completion` and no\n  aggregate or per-check deadline.\n- The candidate.20 catalog is the deterministic source candidate for the\n  `1.1.0` task set. A fresh independent review and seal, complete 17-by-72 calibration,\n  policy-v2 fixed-bank admission, separate complete Official run, publication,\n  and deployment are required. No earlier publication is a fallback.\n- The public catalog contains metadata and commitments, not private task content.\n- Task scores use committed weighted binary checks. A failed hard gate or\n  structural check sets the score to zero; otherwise the evaluator divides\n  passed positive weight by total positive weight. The runner commits only one\n  semantic evaluator result for each sealed response and workspace. A retryable\n  evaluator process failure keeps that evidence pending and reruns only the\n  evaluator on resume. The independent verifier executes the evaluator once and\n  compares the parsed result and exact raw output digest. A failed verifier\n  invocation releases the claim for a later model-free replay. A first\n  successful replay with different output also requires a later confirmation\n  attempt. Publication remains blocked until one exact replay matches.\n- The source-head AIQ measurement contract is `2.0.0`: the Official ranking\n  score is `100 × logistic(theta)` from the admitted fixed Rasch item bank;\n  theta and its conditional Wald interval are reported separately from the raw\n  equal-domain `qualityScore` diagnostic. This contract is not an IQ norm or a\n  150-point scale.\n- Calibration policy `aiq.official-calibration-policy.v2` reports the binary\n  informative-task rate and its 0.50 descriptive target, but does not use that\n  count as a release cliff. Complete semantic coverage, non-uniformity,\n  universal floor and ceiling limits, domain checks, and model and latent\n  spread remain hard gates.\n- Strict pass is strict successes divided by all attributable tasks with a\n  valid semantic task score. Partial scores remain in that denominator; only\n  missing, infrastructure-invalid, runtime-failed, and unscored tasks are\n  excluded. `invalid_tasks` records observed runtime or infrastructure\n  failures, while `missing_tasks` is reserved for an expected cell with no\n  result record. Runtime failures are not semantic zeros. The Wilson interval\n  uses the same sample.\n- The model matrix contains 17 configurations: six Sol, six Terra, and five Luna.\n- The runner performs capability preflight, executes tasks, scores results, and\n  creates signed `aiq.result-package.v4` envelopes.\n- Every result keeps separate runner-observed model and evaluator elapsed time\n  as `latency.wall_ms` and `latency.evaluator_ms`\n  and, when Codex reports it, token usage and a versioned Standard\n  API-equivalent cost estimate.\n- AIQ, Rasch ability, quality, strict pass, ranking, and intervals use only\n  evaluator-backed semantic task scores. Elapsed time, tokens, tool use, and\n  estimated cost are independent efficiency evidence and never change a score.\n- Public evidence labels time as `runner_observed`, provider token source as\n  `provider_reported`, and verifier-checked token and cost evidence as\n  `verifier_recomputed`. Unavailable evidence remains null, not zero.\n- The verifier reconstructs submitted workspaces and replays deterministic\n  evaluators before it signs `aiq.verifier-attestation.v4` evidence.\n- The verifier also provides an offline `diagnose-rescore` audit. It first\n  verifies and replays one source package, then scores the preserved cells with\n  a candidate source, task, evaluator, runtime, and toolchain set. Its\n  create-new report is permanently non-Official and non-ranking. It cannot\n  publish or create an attestation.\n- Production uses three distinct identities: runner, verifier, and publisher.\n- The Web application reads public database views and sends controlled writes\n  through server routes.\n\nThe source-head ordered task-metadata catalog digest is:\n\n```text\nsha256:3580555315d49a62b28b6947491819276dca5b261ade802f10b33808569d1708\n```\n\nIts public source-release digest is\n`sha256:5b651845280ea0b27a8cfc2aec4efe4c149bbbfdeeb0e5a9c883174938b58d69`.\nThe production task-set identity is `aiq-core/1.1.0`; the retained catalog source\nidentity is `aiq-core/1.1.0-candidate.20`. Do not infer any controlled\nidentity from these public digests. The reviewed evaluator identity is\n`sha256:748e0a6c07eb7e3407cc22d50b65eb6d055305cb6e1d719ca3cfd3a109bec809`.\nThe current no-deadline public-safe database task-set identity is\n`sha256:c7481e46c64dbf5ff9f50a85c83608d48390a03cbf9e94a1d89ab36aeb6df89a`,\nand its task-commitment manifest identity is\n`sha256:d8dddd1bc496a1609c3268068fdfdfa4562c589ddfdfec365a6a49caadefe96b`.\nThese are checked-in bindings derived from the reviewed candidate.15 seal.\nProduction activation still requires the exact private corpus, calibration,\nadmission, Official package, and verifier evidence.\nThe checked Core schema\nrequires `runner.identity_kind` to remain `source_only` and\n`runner.built_binary_sha256` to remain null. The shared Rust validator now fails\nclosed on this runner subtree for both Core and Contrast. Contrast does not have\na separate checked-in JSON schema. Each corpus also binds the Node.js and ripgrep\nidentities. The source-only corpus rule and signed per-run runner and complete\nCodex runtime provenance are the executable product contracts. The Codex runtime\nis one private directory that contains exactly the `codex` executable and its\n`codex-code-mode-host` sibling. After the final clean build, the operator retains\na private, unsigned audit receipt with the exact source commit and tree identity\nand SHA-256 values for the native runner, verifier, Codex executable, and Codex\ncode-mode host. The offline native verifier validates this receipt against an\nindependently supplied receipt digest. It is not a database input or published\nartifact. Node.js and ripgrep remain bound by the corpus commitment. Do not infer\na runtime hash from a generated-task tree digest. The accepted AIQ 2.0 publication\nwill be one batch of 17\nconfiguration runs and 1,224 task-level executions.\nElapsed time, provider-token usage, and Standard API-equivalent cost are\nreported separately from AIQ.\n\nThe Web application is a professional analysis workbench. Official evidence\npresents calibrated ability with its conditional 95% interval. Synthetic fixtures\npresent descriptive quality with task-mix sensitivity and never appear as Official.\nScientific context also reports strict pass with a Wilson interval, sample count,\ncoverage, missing cells, runtime state, scoring method, and provenance. It keeps\nsemantic task outcomes separate from runtime, invalid, and missing cells. Cost\nremains an estimated Standard API-equivalent comparison, not an actual ChatGPT or\nCodex subscription bill. Charts use ECharts with SVG rendering and ARIA\ndescriptions. Users can select system, light, or dark color themes. Production\nviews must use only real evidence for the sole production tuple, not synthetic\nor legacy data.\n\n## AIQ Core 1.1.0 source authority\n\nThe `1.1.0` source candidate is candidate.20 at\n`benchmarks/candidates/aiq-core-1.1.0/`. It preserves candidate.19 except for one\nscored, digest-bound final-response reconciliation decision in `tool-use-02`.\nCandidate.19 completed all 1,224 calibration cells but failed policy v2 because\nthe Tool Use domain mean facility was `0.949580`, above the `0.90` ceiling.\nCandidate.20 keeps the policy, 72 tasks, 17-model matrix, supplied tool, and\nworkspace receipt semantics unchanged. It still requires fresh independent\nreview, sealing, calibration, admission, and Official evidence.\n\n`benchmarks/candidates/aiq-core-1.1.0/catalog.json` uses\n`aiq.catalog.v2`. Candidate.20 keeps 71 response contracts and advances only\n`tool-use-02` from workspace-only scoring to a final-response reconciliation\ndecision while retaining its workspace and receipt checks. It keeps the\nseven distinct tool-use constructs. It records 29 private tasks as retained and\n43 as repaired after candidate.15 calibration. It has 72 distinct\nwithin-domain clusters. Every task requires `gold`, `alternate_correct`, `partial`,\n`adversarial_format`, and `empty`; `timeout` is `not_applicable` under natural\ncompletion. The catalog is the sole expected-class authority.\nIts canonical catalog digest is\n`sha256:00e555904daa023f0f7731a0e7e66e12d833641f62eb4aaa59447be6417807d7`.\nIts ordered task-metadata digest is\n`sha256:3580555315d49a62b28b6947491819276dca5b261ade802f10b33808569d1708`.\nIts public release digest is\n`sha256:5b651845280ea0b27a8cfc2aec4efe4c149bbbfdeeb0e5a9c883174938b58d69`.\n\nCandidates.1 through .18 are immutable predecessor evidence. Candidate.5\nremains the durable source for the seven distinct disclosed scenario, operation,\nresult, evaluator, metamorphic, and cross-task substitution contracts. Its\nsource integration was rejected because\nits catalog task-metadata identity was\n`sha256:cfac96630c9efe3153d80ed43effd6e541bef751e1e7f766a52cfb2910fa3fc4`,\nwhile the Rust commitment consumer and public v3 schema still required\n`sha256:393cb2563b2161ccb42dd5a50ea63a7827f4d5c485ca0a98103e80eef3d0fbe6`.\nCandidate.6 removed that duplicate Rust identity authority. Candidate.7 kept the\ncatalog-derived commitment authority and repaired the trusted execution and\nqualification-evidence bridge. Candidate.7 is rejected because its completed-run,\nrecovery, and package paths returned to the active 1.0.7 validator after candidate\npreparation. Candidate.8 carries one provenance-bound validation context through\nthose paths, but its package command derived that context from the saved record and\nits private runtime used Node.js 24.19.0. Candidate.9 requires independent task,\ncorpus, and source inputs for candidate packaging. It binds the complete validated\ncontext through signed-payload serialization and uses Node.js 24.18.0. The active\n1.0.7, Contrast, historical, and Official validators remain\nunchanged. The commitment validator derives the expected candidate identity from\nthe validated embedded catalog and rejects candidate.14 and older identities.\nCandidate.9 is rejected because `debugging-04` declared `src/task.mjs` while its\nprompt, workspace bindings, and weighted evaluator import `src/task.ts`, and\n`instruction-following-05` declared the non-schema field type `undefined` for\n`calculation_note`. Candidate.10 corrects those values to `src/task.ts` and\n`string`, but its location check is rejected because candidate.3's versioned\nresponse contract selected the source locations used to validate candidate.3\nthrough candidate.5. Candidate.11 preserves the corrected leaves and all task,\nevaluator, fixture, and tool semantics. Its separately tracked\n`benchmarks/candidates/aiq-core-1.1.0/task-response-authority.json` file owns the\npublic-safe response mode and locations for every task. The generator validates\nall versioned contracts against that unchanged projection. Candidate.11 is\nrejected because its tracked private validator treated protected\n`expected_file_sha256` inputs as response outputs and inferred final response\nfrom workspace-policy absence. Candidate.12 corrects that existing owner only.\nIt requires one hard-gate `complete_workspace_policy`, detects `response_*`\nchecks for final response, excludes protected inputs from the mutable allowlist,\nand derives workspace locations from progress files plus evaluator `path`\ntargets, with evaluator-source fallback only when both are empty. The immutable\ntasks resolve to 71 workspace responses and one final response: progress is exact\nfor 66 workspace tasks, empty for four, and a strict subset for one.\nCandidate.12 is rejected before model invocation because file loading used lexical\ntask order while candidate validation required checked-catalog order. Candidate.13\nuses the existing catalog-order owner during candidate-only preparation. It also\nselects the v3 commitment schema only on that same candidate preflight route; the\nactive and standalone preflight routes continue to require v2.\nCandidate.13 is rejected operational evidence. Its isolated Jordan qualification\ncompleted all 17 preflight probes and started 88 task cells, but only 56 of 1,224\ncells completed before the five-hour subscription quota reached 100%; seven-day\nusage was 17%. No package, verifier stage, attestation, or qualification artifact\nexists. Candidate.14 keeps every task-facing semantic unchanged and replaces only\nthe quota-infeasible candidate release-qualification shape. Its one 216-cell run\nand package completed, but the real candidate verifier loaded ordinary filenames\nlexically and rejected the catalog-ordered evaluator identity before replay. No\nstage, attestation, or qualification artifact exists. Candidate.15 reuses the\nexisting checked-catalog ordering owner in that candidate-only verifier load path.\nIts complete 1,224-cell calibration had no runtime failures, but policy v2\nrejected 38 universal full-credit tasks, ten universal semantic-zero tasks, only\n22 non-uniform tasks, and six degenerate domains. Candidate.16 keeps policy v2,\nrepairs those task-bank failures, and normalizes only the exact Codex zsh\ntransport before hashing the logical ToolUse command.\nIndependent review rejects candidate.16 because all 43 revised tasks reject a\npublicly declared optional field; six Documentation tasks also score optional\n`next_steps` as required. Candidate.17 repairs that existing evaluator owner and\nthe independently identified 1.1.0/v3/13-view runbook drift.\nIndependent source review rejects candidate.17 because one Security model\nsentence still says 12 public views and the focused regression did not match\nthat wording. Candidate.18 corrects that final documentation owner. Its isolated\nMorgan calibration reached 288 checkpoint results before Codex returned\n`Selected model is at capacity`; task output containing `authentication` caused\nthe adapter to misclassify that temporary capacity event as terminal authentication.\nCandidate.19 routes the exact structured capacity event through existing resumable\nbackpressure and limits provider-failure classification to stderr or structured\n`error` and `turn.failed` messages.\nIts complete calibration then failed only the Tool Use domain facility ceiling;\ncandidate.20 adds the bounded `tool-use-02` response decision described above.\n\nThe exact 42 task-issue closures remain unchanged. The catalog records the\nunauthenticated candidate execution and qualification-evidence bridge as one\nseparate source-integrity closure. It records the candidate.7 end-to-end validation\nfailure as a second source-only closure. Neither closure counts as a task issue.\nThe package-input, Node.js runtime, public response-contract, and private\nresponse-source-owner repairs are four additional source-only closures. None\ncounts as a task issue.\n\nFor each tool-use task, the hard gate requires exactly one total tool call and\none `command_execution` call. It also requires one completed command-line digest for\n`node bin/task-tool.mjs`:\n`sha256:6763cc80f8294b52c6494f1c9891e41a8e3cd1c466ca622377c59643a0466319`.\nThe separate receipt `command_sha256` continues to identify the supplied tool\nfile bytes. The runner retains digest counts only for task-declared required\ncommand identities; exact total and per-tool counts still expose undeclared or\nextra calls. One three-configuration qualification matrix can contain at most\n21 declared digest entries. The runner removes command text from provider\nstdout and stderr evidence, but it\nextracts and preserves the exact semantic final response before that log\nredaction. Candidates.14 through .17 remain inactive and not production-publishable.\nCandidate.18 is the current source candidate; production cutover requires its\nfresh review, seal, complete calibration, admission, and Official evidence.\n\nAIQ Core 1.1.0 sealing also requires one independently supplied\n`aiq.leakage-review.v2` record for every task. Each record binds the reviewer,\nthe reviewer task or thread, review time, source commit and tree, source\nmanifest, task definition, catalog entry, verdict, method, scope, and notes.\nThe sealer copies and hashes the supplied record. It rejects missing, extra,\nrejected, stale, or mismatched records. It never creates a completed review\nfrom task-authored notes. Recorded process separation is review evidence. It is\nnot cryptographic proof that a human reviewer was independent. The v1 review\ncontract remains isolated to frozen 1.0.7 compatibility.\n\nThe shared `aiq-runner` library owns\n`aiq.benchmark-qualification-policy.v2`. Both existing executables use that\nsame in-process implementation:\n\n```sh\ncargo run -p aiq-runner -- qualify-candidate --help\ncargo run -p aiq-verifier -- verify-qualification --help\n```\n\nQualification accepts exactly one predeclared, complete, non-synthetic 3-by-72\nCalibration stage: Sol medium, Terra medium, and Luna medium in that order over\nall 72 catalog-ordered tasks. Its verifier-derived 216 cells must all be\nsemantically complete, and its attestation must come from the predeclared\nverifier. The manifest fixes the candidate, corpus, source, model selection,\nrun, and verifier identities before execution. The final artifact additionally\nbinds the exact signed package, runner, stage, attestation, provenance, and\nmatrix digests.\n\nThis is an end-to-end execution and identity qualification only. It makes no\nprediction-interval, Spearman-correlation, run-variance, or precise-rank claim.\nThose stability-only fields are absent from the v3 manifest and artifact. The\nactive/default production matrix remains all 17 configurations. Any task,\nevaluator, policy, or source-identity revision requires a new reviewed source\nidentity and fresh release evidence.\n\n## Repository map\n\n| Path                 | Purpose                                                                   |\n| -------------------- | ------------------------------------------------------------------------- |\n| `apps/aiq/`          | Scheduled observation orchestration, release validation, and cleanup      |\n| `apps/aiq-runner/`   | Capability checks, task execution, scoring, packaging, and submission     |\n| `apps/aiq-verifier/` | Queue claims, artifact reconstruction, evaluator replay, and attestations |\n| `apps/web/`          | Public Next.js site and controlled server gateways                        |\n| `benchmarks/`        | Public catalog, schemas, and synthetic examples                           |\n| `databases/`         | Desired database state, fresh initializer, and disposable SQL checks      |\n| `openwiki/`          | Architecture, method, operations, and deployment handoff                  |\n\nPrivate tasks, expected outputs, controlled evaluators, signing keys, Codex\nauthentication, and production data must stay outside Git.\n\n## Local synthetic demonstration\n\nUse Node.js `24.15.0` or newer, the npm `11.17.0` version pinned by `package.json`,\nthe stable Rust toolchain selected by `rust-toolchain.toml`, and the locked\ndependencies. `cargo make fmt`\nalso requires a separately managed nightly rustfmt toolchain.\n\n```sh\nnpm ci --ignore-scripts\ncargo run -p aiq-runner -- demo\nnpm run dev\n```\n\nOpen `http://localhost:3000`. When both public Supabase variables are absent in\ndevelopment, the site uses checked-in synthetic data. Production fails closed\nwhen its configuration is incomplete.\n\nUseful runner commands:\n\n```sh\ncargo run -p aiq-runner -- matrix\ncargo run -p aiq-runner -- validate --public-tasks benchmarks/examples/tasks\ncargo run -p aiq-runner -- validate-core-corpus --help\ncargo run -p aiq-runner -- validate-contrast-corpus --help\ncargo run -p aiq-runner -- --help\ncargo run -p aiq-verifier -- --help\ncargo run -p aiq-verifier -- diagnose-rescore --help\n```\n\n## Validation\n\nInstall the Playwright browsers once on a fresh host, then run the complete local\nbrowser gate:\n\n```sh\nnpm exec --workspace @aiq/web -- \\\n  playwright install --with-deps chromium firefox webkit\ncargo make check\ncargo make verify\n```\n\n`check` is the complete read-only source gate. It checks the database schema,\nTypeScript, Rust and TypeScript lint rules, vstyle, and all Rust and TypeScript\ntests. `verify` extends `check` with one Web build and every local browser\nacceptance suite. `fmt` is an independent mutating action; it is not a dependency\nof either gate. Native release builds, production browser checks, database\nruntime checks, deployment, and publication remain separate contracts. Do not run\ncomponent tasks again in the same validation pass. Coverage instrumentation is\nopt in with `cargo make test-typescript-coverage`.\n\nThe two subscription smokes are ignored and opt in. Each consumes one Codex\nsubscription attempt.\n\n```sh\ncargo make smoke-subscription\ncargo make smoke-controlled-subscription\n```\n\nThe public-task smoke validates a fixed example. The controlled-task smoke needs\noperator-supplied private task, evaluator, corpus, runtime, workspace, and Codex\ninputs. Neither smoke creates a benchmark result.\n\n## Database initialization\n\n`databases/schema.sql` is the sole desired database state.\n`databases/init.ts` is the only production initialization entry point. There is\nno migration chain. It opens\none PostgreSQL connection and applies the schema plus public reference data in\none transaction. It accepts the direct host or exact port-5432 session pooler\nidentity for personal Supabase project `xxnszykaeapolqdnhalx`. An explicit\ntest/development override accepts\nonly a loopback target and cannot apply in production. It rejects a database\nthat already contains the AIQ schema, gateway roles, or either exact AIQ\nStorage bucket identity. Apply this one greenfield desired state to the existing\ntarget project only after its AIQ namespace is empty. If AIQ residue exists, the\noperator must\nremove only `aiq_private`, the two AIQ gateway roles, and the exact AIQ-owned\npublic views and RPC overloads. Preserve all Supabase-managed and non-AIQ\nobjects. This cleanup is a deployment prerequisite, not a migration or\ncompatibility path. The schema creates the `aiq-submission-packages` and\n`aiq-runner-artifacts` Storage buckets as private. The preflight rejects either\nexisting bucket identity. Do not create the buckets in a separate operator step.\nThe preflight enumerates the 13 canonical public view names and all public RPC\nnames from the desired state. It rejects every overload of those exact RPC\nnames without matching unrelated public objects.\n\n```sh\nAIQ_DATABASE_URL='<direct-or-session-pooler-url>' \\\nAIQ_PRODUCTION_REFERENCE=/controlled/production-reference.json \\\ncargo make init-database\n```\n\nFor an empty AIQ namespace, the production reference must contain the real\ncontrolled, non-synthetic AIQ Core `1.1.0` v3 corpus commitment, its real canonical\n`published_at` timestamp, and\nexactly three public identities: runner, verifier, and publisher. Prepare it\nonly after the controlled corpus passes model-free validation, the operator\nverifies the final native build, and one real signed non-synthetic 17-by-72\npackage passes native verifier replay; the repository contains no substitute\nproduction reference. Retain the private final-build audit receipt separately.\nDatabase initialization does not accept or validate that receipt.\nA successful initialization receipt must report aggregate scoring `1.0.8`, both public\ncatalog identities, 72 tasks, 17 model configurations, and three nodes.\n\nUse one initialized disposable database for production-shape smoke and\ncalibration publication checks:\n\n```sh\ncargo make smoke-database\nAIQ_DATABASE_URL='<direct-or-session-pooler-url>' cargo make smoke-calibration-database\n```\n\nUse a separate fresh PostgreSQL 17 database for the deterministic synthetic\nflow:\n\n```sh\npsql \"$AIQ_DATABASE_URL\" -X --set ON_ERROR_STOP=1 \\\n  --file databases/schema.sql\npsql \"$AIQ_DATABASE_URL\" -X --set ON_ERROR_STOP=1 \\\n  --file databases/synthetic-demo.sql\npsql \"$AIQ_DATABASE_URL\" -X --set ON_ERROR_STOP=1 \\\n  --file databases/integration.sql\n```\n\nDo not apply the synthetic flow to the initialized production-shape database\nor to production.\n\n## Production data flow\n\n1. The runner validates the controlled corpus, toolchain, and capability\n   manifest.\n2. It executes the selected tasks and runs each formal evaluator against the\n   sealed model evidence. A retryable evaluator process failure keeps the model\n   evidence pending for evaluator-only recovery. The runner writes\n   content-addressed artifacts.\n3. It scores the run, records efficiency evidence, and signs one v4 result\n   package.\n4. `POST /api/submissions` stores the exact package bytes and queues the package\n   as unverified.\n5. The verifier claims the package, reconstructs the workspaces, and executes\n   each deterministic evaluator once for that claim attempt. An operational\n   replay failure releases the claim for retry without changing the retained\n   package or invoking a model.\n6. `POST /api/verifications` stages the normalized batch and records the signed\n   verifier attestation.\n7. A distinct publisher identity completes publication through the gateway.\n8. Public security-invoker views supply the Web application.\n\nOfficial means a complete, non-synthetic 17-by-72 run with valid task-set\n`1.1.0`, task-scorer `1.0.6`, aggregate-scorer `1.0.8`, and measurement `2.0.0`\nbindings that completed this flow and was published as\n`trusted_verified`. A complete synthetic fixture uses the\n`synthetic_complete` classification, has no Official AIQ value, and is never\nranking eligible. There is one submission, native verification, and publication\npath.\n\n## Official paid-work boundary\n\nThe only Official execution and publication path runs `aiq-runner` and\n`aiq-verifier` natively on the controlled Apple Silicon macOS host with direct\nnetwork access. Use the release binaries in this order:\n`admit-permissions`, `preflight`, `run`, `score`, `package`, `submit`, and then\nverifier replay. `admit-permissions` is model-free; `preflight` is the first paid\nstep. Only its exact configuration probes and runnable task cells in `run`\ninvoke models. Scoring, packaging, submission, verifier replay, and publication\ndo not invoke models. The same private admission receipt binds preflight through\npackage.\nProvide the runner signing key only to `package`, the submission token only to\n`submit`, and verifier credentials only to the verifier command.\nThe source runner targets native macOS. Linux and Docker remain future\ndeployment targets. A frozen `aiq` release built from this source starts the\nobservation scheduler at `03:00` and `15:00` UTC. It selects one canonical\n12-hour slot, holds a global nonblocking lock, and does not start a second\nscheduler. An installed frozen release keeps its existing behavior until an\noperator replaces it. A self-contained release\nstores the pinned runner and verifier binaries and an exact Git source bundle.\n`aiq` restores the clean detached source at a stable per-slot path below the\nprivate `state_root/scratch` directory. This path stays outside the macOS\nplatform-minimal roots that model processes can read. The command does not use a\nrepository worktree at run time. On a host fixed to the `America/New_York` time\nzone, the macOS `launchd` template wakes at 11:05, 11:35, 23:05, and 23:35 local\ntime. The four wakes cover EST and EDT with one bounded retry for each UTC slot.\nOfficial task dispatch must begin during the first two hours of a slot, and the\nv2 configuration requires all 32 supported workers for the fixed 1,224-cell\nmatrix. A late wake does not start a new matrix. Temporary selected-model capacity and\nsubscription quota, usage, or rate limits are persisted as non-terminal backpressure:\ncompleted cells stay in the same checkpoint, rejected cells remain pending, and later\nscheduled wakes resume the oldest blocked slot before they can start newer paid work. The\nscheduler starts Official and Speed as sibling publication paths for the same\nslot. Official keeps its two-hour model-dispatch grace. Speed has an independent\n12-hour slot window. After the scheduler grants dispatch, neither path waits for\nthe other path. A slow or failed path cannot block the other path's dispatch or\npublication. Each path writes retained status below its own slot directory, and\n`aiq status` composes both outcomes. A completed run\nwith a non-semantic infrastructure result is retained as unpublished evidence.\nIt is not retried or presented as an AIQ score. Provider-capacity backpressure is not a\ncompleted result and is therefore the sole exception to that terminal rule.\nThe subscription runner uses a protected copy of `~/.codex/auth.json` in an\nisolated per-release `CODEX_HOME`; it does not reuse the interactive Codex home\nas its writable runtime directory. It also uses a private two-file copy of the\nChatGPT app's `codex` and `codex-code-mode-host` executables. Capability\npreflight succeeds only after Codex completes one command and writes the exact\ncontent-bound marker in a fresh disposable workspace.\nSee [Operations and Validation](openwiki/operations.md) for the native command\ncontract. Repository support does not prove that private inputs, credentials,\nor live model capabilities are configured.\n\n## Continuous observations\n\nNormal/Fast transport measurements are auxiliary evidence. `observe-speed`\nreads the live Codex model catalog before any paid turn, records an exact\navailable, unsupported, or unavailable state for each selected configuration,\nand runs paired Normal/Fast fixed-response trials only for advertised modes.\nIt records completion, total elapsed time, aggregate output throughput, token\nusage, tool use, and estimated ChatGPT credits. It does not calculate or modify\nAIQ. The current Codex JSONL stream does not expose a trustworthy first-token\ntimestamp, so TTFT and post-first-token throughput remain explicit unavailable\nvalues instead of estimates.\n\n```sh\ncargo run -p aiq-runner -- observe-speed --help\ncargo run -p aiq-runner -- submit-speed --help\ncargo install --locked --path apps/aiq\naiq status --config /absolute/private/path/to/continuous-observation.json\naiq doctor --config /absolute/private/path/to/continuous-observation.json\naiq run --config /absolute/private/path/to/continuous-observation.json\naiq run --config /absolute/private/path/to/continuous-observation.json \\\n  --slot 2026-08-12T03-00Z\n```\n\nUse `--slot` only for one known canonical UTC slot. Official task dispatch can\nstart only during the first two hours of its current slot. The frozen runner\ncan resume an unchanged checkpoint during the same slot only when it contains\nno indeterminate in-flight cell. A checkpoint with explicit subscription\nbackpressure can also resume after the dispatch grace or 12-hour slot window;\nit reuses the exact admitted preflight and never replaces completed cells. Other\ncheckpoints with sealed pending evaluator work can also resume that work after\nthe window without another task-model invocation. This rule includes a retryable\nevaluator process failure. An indeterminate model cell\nstill fails closed after all sealed pending evaluator work is recovered. Other\nlate slots can continue only when the complete Official run output already\nexists and only scoring or publication remains. `aiq` recognizes that output\nonly as an `aiq.run.v4` document with all 1,224 results; the runner's\ncreate-once reservation is not a completed run. Otherwise, `aiq`\nrecords a terminal missed or unpublished state without new model work. Speed\nmodel dispatch can start during its own 12-hour slot. An existing Speed batch\ncan resume submission after that window without new model work. A\nterminal slot remains a no-op. If another\nobservation owns the global lock, a scheduled `run` coalesces successfully\nwithout starting another model process; `doctor` reports the contention instead.\n\nEach runner, verifier, and evaluator step runs below an internal supervisor that\nowns a separate process session. Runner-created model and evaluator process\ngroups remain in that session. A private pipe binds the supervisor to the\nuser-facing `aiq` parent. If that parent exits or is killed, the pipe closes and\nthe supervisor repeatedly sends `SIGTERM`, then `SIGKILL`, to every remaining\nsession process before it exits. This no-orphan boundary does not depend on\n`launchd` process-group cleanup.\n\nStart from `config/continuous-observation.example.json` and\n`config/com.acgbox.aiq.continuous-observations.plist.example`. Keep the concrete\nconfiguration and `launchd` plist outside Git. Use `aiq install-release` once to\ncopy the minimal frozen release, create its source bundle, and print the release\nmanifest digest. Install the release in a versioned directory outside the\nrepository. The private v2 configuration contains stable runtime paths, limits,\nthe endpoint, the manifest digest, and optional non-secret unattended provider\nmetadata. It does not contain a source worktree path, worker executable path,\nprovider credential, or consumer secret.\n\n`cargo install` is sufficient for local operator use. An unattended service\nmust pin `apps/aiq/package.nix` in the host configuration so an unrelated Cargo\ninstall cannot replace the scheduled executable.\n\nSet `official_jobs` to `32`; lower values are rejected before model work. The\n`aiq run` accepts either all four explicitly supplied consumer variables or no\nconsumer variables. Partial ambient delivery fails closed. When all four are\nabsent, `aiq` requires the complete `unattended_secrets` metadata, reads the\nexact Keychain bootstrap, performs one Universal Auth login, and retrieves only\nthe four fixed `prod:/aiq` keys. It removes the provider session before it starts\na downstream step. Provider credentials and tokens do not reach workers. The\norchestrator gives the signing key only to `package`, the submission token only\nto submission steps, and verifier credentials only to the verifier. Each owner\nuses a fresh isolated `CODEX_HOME` directory for the slot. A retryable slot\nretains checkpoints and raw artifacts. Checkpoint v10 distinguishes\nindeterminate model work from sealed pending evaluator work. The latter resumes\nfrom the same model response and workspace without another model invocation. A\nretryable evaluator process failure stays in this pending state and cannot\ncreate a terminal run. Provider-declared temporary model capacity and subscription limits\nleave the affected cells pending under `aiq.subscription-backpressure.v1`; the runner\nalso migrates v9 checkpoints to the v10 evaluator-resume shape and legacy v8\ncheckpoints that incorrectly committed those limits as terminal results. On resume, `aiq` revalidates the permission\nadmission, complete Official run, submission receipts, and verifier receipt\nbefore it reuses them. It stores non-success verifier records in a private\nappend-only attempt log. A create-once success receipt is valid only when its\npackage SHA-256 and idempotency identity match the exact local package. Copied\ncredentials are removed after each invocation. A terminal slot keeps both\nowner-status records, the compact batch,\npackage, score, attestation, and receipts. It removes the detached source, raw\nlocal artifacts, replay scratch, checkpoints, and disposable workspaces.\n\n`launchd` invokes the pinned `aiq run --config ...` command directly. Use\nabsolute AIQ and configuration paths, supply `HOME`, `USER`, `LOGNAME`, and the\npinned execution `PATH`, and do not set a repository working directory. The\nprovider identity must grant only the four fixed source keys. The runtime keeps\nthe Keychain bootstrap and short-lived provider token inside AIQ.\n\nThe already-provisioned provider target is external frozen state. Do not run\nsetup to reconcile, rotate, or replace it. For a new exact target only, the\nhidden `aiq operator provision-unattended --config ...` command uses\n`config/unattended-provider-provision.example.json`. It refuses an existing\nKeychain account or provider identity, creates only the fixed identity,\nfour-key privilege, Universal Auth method, and Keychain bootstrap, and rolls\nback only known intermediate writes.\n\n## Security boundaries\n\n- Keep runner, verifier, and publisher credentials separate.\n- Keep privileged Supabase values in server-only environment variables.\n- Keep `aiq-submission-packages` and `aiq-runner-artifacts` private.\n- Use RLS and the narrow database RPCs; do not write private tables from the\n  browser.\n- Put authentication, request limits, and a WAF in front of write routes.\n- Run the Storage reconciliation worker before the deletion worker.\n- Treat readiness responses as bounded dependency evidence, not deployment\n  proof.\n\nSee [OpenWiki quickstart](openwiki/quickstart.md),\n[operations](openwiki/operations.md), and\n[deployment handoff](openwiki/deployment-handoff.md) for the maintained details.\n\n`aiq.wiki` is canonical, and `www.aiq.wiki` returns a permanent `308` redirect\nthat preserves the request path. Automatic Vercel project and branch aliases can\nbe removed only transiently because a later deployment can recreate or reassign\nthem. A deployment-specific URL is intrinsic to its retained deployment. The\ncurrent generated Vercel surfaces emit `noindex`.\n",
  "bytes": 39431,
  "sha": "f1d2743f7a340d0d8ae2fa3c46949193a98a81a51cd9421ce360bcc409e581e9",
  "repo_slug": "acg-box/aiq",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/okf_acg_box_aiq_openwiki_index_md_472b93ca/readme"
}