JOSSeph
JOSSeph (Java OSS metrics Software) is a Docker-first reproducible pipeline for extracting metrics from GitHub Java repositories.
One run does this:
- Load and validate config.
- Build extractor runtime.
- Process repositories in parallel.
- Write per-tool artifacts.
- Write one run summary.
The useful part is not "metrics in general". The useful part is a strict, reproducible contract:
- input is one YAML file plus one repository list
- the primary execution target is
docker compose run --rm josseph <config> - the Docker runtime makes execution platform-independent
- output is a predictable directory tree under
results/ - every partial failure is supposed to end up in
summary.json - a cached result is reused only when both
<tool>.parquetand<tool>.jsonexist (and, for revision-bound extractors with a pinned commit, the recorded commit hash matches)
What a run produces
For repository https://github.com/example/project.git and tools github and
ck, the steady-state output is:
results/
example@project/
github.parquet
github.json
ck.parquet
ck.json
runs/
20260328T120000Z/
summary.json
github.json has this excerpted shape:
{
"commit_hash": "",
"requested_commit_hash": null,
"metric_binding": "observation-bound",
"collected_at_utc": "2026-03-22T12:34:56Z"
}
summary.json is the run-level source of truth. Excerpt:
{
"run_id": "20260322T150000Z",
"status": "success",
"started_at_utc": "2026-03-22T15:00:00Z",
"finished_at_utc": "2026-03-22T15:00:12Z",
"duration_seconds": 12.0,
"exit_code": 0,
"summary": {
"repository_count": 1,
"affected_repository_count": 0,
"repository_failure_count": 0,
"extractor_failure_count": 0,
"failed_run_count": 0,
"skipped_run_count": 0
},
"repository_failures": [],
"extractor_failures": [],
"failed_runs": [],
"skipped_runs": []
}
If you only read one file after a run, read results/runs/<run-id>/summary.json.
It tells you which repository failed, which extractor failed, which results were
skipped as cached, and whether the process exited 0, 1, or 2.
Start here
- Getting Started for the fastest real run.
- Config for the exact YAML contract.
- Output for artifact schemas and failure cases.
- Examples for concrete success and partial-failure runs.