GET/resultsPoll status and results
Clinician review is done by people, so results are not instant. Poll this endpoint.status moves through received, in_review, delivered, and total / done show progress. received means no clinician has started yet; it becomes in_review as soon as any item is being reviewed. Only delivered includes the report and items.
curl "$BASE/results?project_id=YOUR_PROJECT_ID" \
-H "Authorization: Bearer $API_KEY"While in review
{ "ok": true, "project_id": "...", "status": "in_review", "total": 200, "done": 142 }When delivered
{
"ok": true,
"project_id": "...",
"status": "delivered",
"total": 200,
"done": 200,
"report": {
"accuracy": { "value": 0.8, "correct": 160, "assessable": 200, "basis": "verdict+class" },
"critical_misses": [ ... ],
"per_class": { ... },
"qa": { "mean_agreement": 0.86, "reviewed_items": 200, "disagreements": 12 }
},
"items": [
{ "idx": 0, "content": { "case_id": "case_001", ... },
"label": { "verdict": "Correct", ... }, "labeled_at": "..." }
]
}Each item is reviewed by multiple licensed clinicians, and the qa block reports their mean agreement and how many items needed adjudication — so you can trust the numbers.
The assurance block states who stands behind the result: licensed clinicians, independent of you and of the system under evaluation, with Senebiclabs accountable for the findings. Their identities are never disclosed.
Scoring contract: to get the accuracy report, items must carry aprediction and your fields must use these exact names: verdict (Correct / Incorrect / Partial), correct_label (the corrected class for a wrong verdict), and critical_miss (a structured field that populates the report’s critical misses). A wrong verdict with no correct_label is excluded, never guessed. label and create projects skip scoring and return every reviewed item in items as a content-and-label pair.
Free-text evaluations: when the model’s output is prose (an agent’s answer, a RAG response) and the task has no correct_label field, as in grounding_eval, reasoning_eval and triage_eval, there is no class to correct to. accuracy.value is then the pass rate, the share of cases clinicians judged Correct, and every failure counts. accuracy.basis says which applies: verdict for free text, verdict+class when a corrected class is required. Free-text reports have no confusion matrix.