Failure libraryMemory across versions
The endpoints so far evaluate one version. These give the evaluation a memory: what your model always gets wrong, whether a fix held, and a permanent suite every future release is tested against. Tag a run with model_version on POST /projects.
Record a finished evaluation
curl -X POST "$BASE/failures/capture" \
-H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
-d '{ "project_id": "...", "model_version": "v2.3" }'
{ "ok": true, "recorded": 12, "fixed": 4, "regressed": 1 }Failures are upserted by your case id, so a case that keeps failing accumulates occurrences rather than duplicating. Cases this run passed that the library holds open are closed as fixed, tagged with the version that fixed them. A case that was fixed and fails again becomes regressed, never plain open — a fix that did not hold is the most important thing the library knows.
Query and aggregate
GET $BASE/failures?severity=Critical&clinical_domain=Respiratory&status=open
GET $BASE/failures/patterns
{ "patterns": {
"by_status": { "open": 61, "fixed": 22, "regressed": 4 },
"by_error_category": { "Missed red flag": 19, "Triage failure": 14 },
"serious_by_domain": { "Respiratory": 12, "Cardiac": 9 },
"recurring": [ { "case_key": "PMX-47", "occurrences": 3, "status": "regressed" } ]
}}Promote, then re-run
POST $BASE/benchmark/promote { "min_severity": "High" }
POST $BASE/benchmark/run { "model_version": "v2.4" }
{ "ok": true, "project_id": "...", "cases": 34, "compare_with": "<id>,<id>",
"runs": [ { "project_id": "...", "cases": 34, "compare_with": "<id>,<id>" } ] }benchmark/run seeds a new evaluation with every benchmark case, replaying each stored input verbatim — a regression test is only a test if the input does not drift between runs. Cases are graded by the same task config that found them. Cases from several projects graded the same way go into one run, and compare_with lists all of those projects, ready to pass to /compare. If your benchmark mixes tasks graded differently (a triage eval and a grounding eval, say), you get one run per task in runs, each with its own compare_with, and the top-level project_id is omitted.
The release gate, end to end
POST /benchmark/run { "model_version": "v2.4" } -> project_id, compare_with
GET /results?project_id=... -> poll until delivered
GET /compare?baseline=<compare_with>&candidate=<project_id>
-> verdict: block | review | pass
(with several runs: one /compare per entry in `runs`)
POST /failures/capture { "project_id": "...", "model_version": "v2.4" }Four calls. The last folds the new results back into the library, so the next release is tested against everything learned so far.