SenebiclabsAPI referenceWebsiteGet an API key →
Endpoints

POST/projectsCreate a project

Create a project and get back a project_id to push items to.

Start from a template (recommended)

Pick what you want to achieve and we build the project for you — no config to author. List the outcomes with GET /templates:

  • model_evaluation — grade your model’s outputs → accuracy + safety scorecard
  • data_labeling — your data back, labelled → labelled dataset + summary
  • rlhf_preference — pick the better of two responses → preference pairs for RLHF
  • gold_answers — write the ideal answer → gold dataset for fine-tuning
  • case_review — judge whether AI helped or hurt on full cases → audit dataset + impact distribution
  • benchmark_creation — author challenging test cases → an evaluation benchmark
  • rubric_creation — design the scorecard your model is graded against → a reusable grading rubric
  • contradiction_creation — plant a clinically important conflict between a record and a source → a contradiction test set with the expected safe behaviour
  • adversarial_prompts — write probes that expose model gaps → a red-teaming test set
  • fact_checking — highlight errors in an answer, rewrite it, cite a source → accuracy + a corrections dataset
  • dialogue_creation — author realistic patient-clinician dialogues → synthetic training data
  • response_ranking — rank two answers on accuracy/empathy/clarity/safety → preference pairs with per-axis scores
  • clinical_safety_eval — full safety review: correctness, triage, red flags, reasoning → safety scorecard with severity + failure modes
  • triage_eval — is the urgency right → triage accuracy split into under-triage, over-triage, missed emergencies
  • reasoning_eval — score the reasoning step by step → where in the reasoning it fails
  • grounding_eval — retrieval, grounding, citations, hallucination, conflicting evidence → RAG failure breakdown
  • agent_trace_eval — grade a multi-step agent’s whole trajectory → where it first went wrong, which steps failed and how

Create from one, supplying your own classes (label set) where it applies:

curl -X POST "$BASE/projects" \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{
    "name": "Triage model eval",
    "template": "model_evaluation",
    "classes": ["Routine", "Urgent", "Emergency"],
    "webhook_url": "https://your-app.com/hooks/senebiclabs"
  }'

That is all most projects need. The rest of this section is the advanced path — authoring a full config yourself.

Second reading. Authoring templates (gold_answers, benchmark_creation, contradiction_creation, adversarial_prompts, dialogue_creation) are written by one clinician, so a second, different clinician reads every item before it counts. The reader approves it, edits and approves it, or sends it back with a reason, and it is then written again. Only approved items are delivered. Each one carries "second_reading": {"approved": true, "rounds": 1, "edited_by_reader": false, "by_senior_reviewer": false} in your results, where rounds counts the drafts it took, and the report’s second_reading block totals these. It is on by default for authoring projects; pass "second_reading": false to POST /projects to turn it off. Judgment templates don’t use it: several clinicians review each item instead.

Agent traces. For agent_trace_eval, send trace as a list of steps (objects or strings) or as text. Clinicians see it as numbered steps; your results keep it exactly as you sent it. The report adds clinical.steps: where trajectories first break (median_first_failed_step and the distribution).

Tune a template to your own rubric. GET /templates also returns each template’s full eval_config. Take the closest one, edit it to fit your exact task (add rating axes, change fields or context), and submit it as a custom eval_config below instead of template — so you start from a working, validated config, not a blank page.

Custom config (advanced)

Define your own task from scratch:

curl -X POST "$BASE/projects" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Clinical response evaluation",
    "eval_config": {
      "title": "Clinical response review",
      "purpose": "evaluate",
      "schema": {
        "input": "text",
        "context": [
          { "key": "scenario",   "label": "Patient message" },
          { "key": "prediction", "label": "Model response" }
        ],
        "classes": ["Routine", "Urgent", "Emergency"],
        "case_id_field": "case_id",
        "fields": {
          "verdict":       { "type": "single", "options": ["Correct", "Incorrect", "Partial"], "required": true },
          "correct_label": { "type": "from_classes", "visible_when": "verdict!=Correct" },
          "critical_miss": { "type": "structured" },
          "notes":         { "type": "text" }
        }
      }
    },
    "webhook_url": "https://your-app.com/hooks/senebiclabs"
  }'

Response

{
  "ok": true,
  "project_id": "fc64fb22-...",
  "webhook_secret": "a28e0736cb92..."
}

Save the webhook_secret. It is returned once, only when you register a webhook_url, and is used to verify webhook authenticity (see Webhooks). Treat it like a password.

This example is an evaluation project: each item carries a prediction, clinicians return a verdict of Correct, Incorrect, or Partial, and the report scores accuracy. For a creation project, omit prediction and setfields to the labels you want produced; the results come back as content-and-label pairs with no scorecard.

Questions? senebiclabs@gmail.com