SenebiclabsAPI referenceWebsiteGet an API key →
Endpoints

GET/compareRegression between two versions

Evaluate a new model version against the same cases, then diff the two runs. Cases are matched on your case id, so the candidate run only needs to carry the same ids as the baseline.

curl "$BASE/compare?baseline=PROJECT_A&candidate=PROJECT_B" \
  -H "Authorization: Bearer $API_KEY"

If a run was split across several projects (one per specialty, say), pass them all, comma-separated: baseline=<id>,<id>&candidate=<id>. Each side is pooled. A case id must be unique within a side; one that appears in two projects on the same side is refused with a 422 naming it, because pooling would silently keep only one of them.

{ "ok": true, "comparison": {
  "matched": 4,
  "pass_rate": { "baseline": 0.5, "candidate": 0.75, "delta": 0.25 },
  "fixed":     { "count": 2, "cases": [...] },
  "regressed": { "count": 1, "cases": [{ "case_id": "PMX-2", "severity": "Critical" }] },
  "clinical_deltas": { "missed_emergency": { "baseline": 1, "candidate": 0, "delta": -1 } },
  "verdict": { "recommendation": "block", "serious_regressions": 1,
               "still_failing_serious": 0,
               "reason": "1 case(s) that passed before now fail at Critical/High severity" }
}}

fixed and regressed are never netted off against each other. In the example above every headline metric improved and the verdict is still block, because one case that used to pass now fails critically — which is the entire reason to keep a regression suite. recommendation is block (a serious regression), review (a regression), or pass. It judges what the release changed. Known failures that are still failing do not move it, but they are never hidden: still_failing_serious counts those still failing at Critical/High severity, and reason names them. A pass with still_failing_serious above zero means “nothing new broke”, not “safe to ship”.

Cases present in only one run are listed in only_in_baseline / only_in_candidate and excluded from the comparison, so a benchmark that quietly drops a case cannot manufacture a clean scorecard.

Questions? senebiclabs@gmail.com