SenebiclabsAPI referenceWebsiteGet an API key →
Endpoints

GET/resultsPoll status and results

Clinician review is done by people, so results are not instant. Poll this endpoint.status moves through received, in_review, delivered, and total / done show progress. received means no clinician has started yet; it becomes in_review as soon as any item is being reviewed. Only delivered includes the report and items.

curl "$BASE/results?project_id=YOUR_PROJECT_ID" \
  -H "Authorization: Bearer $API_KEY"

While in review

{ "ok": true, "project_id": "...", "status": "in_review", "total": 200, "done": 142 }

When delivered

{
  "ok": true,
  "project_id": "...",
  "status": "delivered",
  "total": 200,
  "done": 200,
  "report": {
    "accuracy": { "value": 0.8, "correct": 160, "assessable": 200, "basis": "verdict+class" },
    "critical_misses": [ ... ],
    "per_class": { ... },
    "qa": { "mean_agreement": 0.86, "reviewed_items": 200, "disagreements": 12 }
  },
  "items": [
    { "idx": 0, "content": { "case_id": "case_001", ... },
      "label": { "verdict": "Correct", ... }, "labeled_at": "..." }
  ]
}

Each item is reviewed by multiple licensed clinicians, and the qa block reports their mean agreement and how many items needed adjudication — so you can trust the numbers.

The assurance block states who stands behind the result: licensed clinicians, independent of you and of the system under evaluation, with Senebiclabs accountable for the findings. Their identities are never disclosed.

Scoring contract: to get the accuracy report, items must carry aprediction and your fields must use these exact names: verdict (Correct / Incorrect / Partial), correct_label (the corrected class for a wrong verdict), and critical_miss (a structured field that populates the report’s critical misses). A wrong verdict with no correct_label is excluded, never guessed. label and create projects skip scoring and return every reviewed item in items as a content-and-label pair.

Free-text evaluations: when the model’s output is prose (an agent’s answer, a RAG response) and the task has no correct_label field, as in grounding_eval, reasoning_eval and triage_eval, there is no class to correct to. accuracy.value is then the pass rate, the share of cases clinicians judged Correct, and every failure counts. accuracy.basis says which applies: verdict for free text, verdict+class when a corrected class is required. Free-text reports have no confusion matrix.

Questions? senebiclabs@gmail.com