Job status
An open file format any pipeline can write, so the catalog can show what is fresh, failing, stuck or stale.
Every job that produces data for the catalog (an Airflow DAG, a GitHub Actions workflow, a cron script, a crawler) writes a small JSON file saying how its last run went. The explorer reads those files to show freshness and failures, both on a Status page and as a badge on each dataset.
- Any orchestrator can write it. It is a JSON file in a bucket or folder. No API, no agent, no database.
- The site stays static. The browser reads the files directly, so status is live without rebuilding the site.
- One write per run is enough. Writing at start (
running) and at the end makes “stuck” detection possible, but is optional.
Layout
Under a status prefix (configured in status.sources of dataherb.config.yml), each job owns a folder named after its id:
<prefix>/<job_id>/latest.json current state, overwritten on every write
<prefix>/<job_id>/runs/<YYYYMMDDTHHMMSSZ>-<run_id>.json one file per run (history)
latest.json is what the site and dataherb status check read. Run files are named with the start time first, so a lexical sort is a time sort.
The document
{
"spec": "dataherb.status/v1",
"job": {
"id": "sales-export",
"name": "Nightly sales export",
"owner": "commercial-analytics@example.com",
"orchestrator": "airflow",
"expected_interval": "P1D",
"max_duration": "PT2H"
},
"run": {
"id": "scheduled__2026-10-07T02:00:00+00:00",
"status": "success",
"started_at": "2026-10-07T02:00:12Z",
"finished_at": "2026-10-07T02:14:39Z",
"message": "32 files processed"
},
"last_success": { "id": "...", "status": "success", "started_at": "...", "finished_at": "..." },
"datasets": [{ "id": "demo-daily-sales", "rows": 11680 }],
"checks": [{ "name": "no_null_region", "status": "fail", "message": "12 rows" }]
}
run.statusis one ofqueued,running,success,partial,failed,skipped,cancelled.last_successis carried forward by writers, so a reader knows a job is stale even when its latest run failed.dataherb status emitdoes this for you.datasets[].idlinks the job to catalog datasets; datasets can also name their job withstatus_jobin their metadata.- Any failing
checks[]entry makes a successful run degraded.
Health
Readers derive one health value per job from latest.json and the current time. The first rule that matches wins.
| Health | When |
|---|---|
unknown |
no run recorded |
stuck |
running or queued for longer than max_duration |
stale |
running or queued, and the last success is older than expected_interval × 1.5 |
running |
running or queued |
failing |
the run failed or was cancelled |
stale |
last success older than expected_interval × 1.5, or never succeeded |
degraded |
the run is partial, or a check failed |
healthy |
the run succeeded or was skipped |
The grace factor is configurable (stale_grace, default 0.5), so a daily job turns stale 36 hours after its last success. A dataset shows the worst health of its jobs.
Writing status files
The dataherb CLI writes both files, carries last_success forward and sets Cache-Control: no-cache:
dataherb status emit --target s3://bucket/_dataherb/status/ \
--job-id sales-export --status running --expected-interval P1D --max-duration PT2H
# ... the job runs ...
dataherb status emit --target s3://bucket/_dataherb/status/ \
--job-id sales-export --status success --message "32 files processed"
Ready-made emitters: Airflow callback, GitHub Actions (via the explorer’s emit-status action), and plain shell.
To alert outside the site, run dataherb status check on a schedule. It prints every job’s health and exits 1 when any job is failing, stuck or stale.
The full specification, with every field, is docs/job-status-spec.md; the JSON Schema is job-status.schema.json.