← Overlap Desk / API
Tokens

Drive Overlap Desk from your own code

A reading, not a verdict on your biology. The model reads the overlap analysis your browser (or your script) computed; it never recomputes a number, and it is never sent your intervals. Overlapping intervals show co-location, not that one set regulates the other.

Everything the web page does is available over HTTP. Analyse your two BED files with the page's own overlapkit.js (gtars 0.9.2 semantics - merged bases, Jaccard, coverage both ways, overlap coefficient, intervals with a hit, nearest gaps, a seeded shuffle test and a universe Fisher test - checked against gtars 0.9.2, bedtools 2.26.0 and scipy), send the facts, and get back a verdict (sound, caveated, unreliable) and either a reading of every metric and chance test or a gtars script that reproduces every number and applies the fixes. The natural loop: analyse, read, script, fix the sets, re-check.

Two lanes: the task field

taskwhat you getextra input
interpretA reading of each of the seven metrics (Jaccard, both coverages, overlap coefficient, intervals with an overlap each way, nearest gap), each chance test with the browser's result (enriched, depleted, not significant, not run), your claims judged against the facts, and what the overlap cannot show.none
scriptThe fixes (renaming contigs, dropping invalid or out-of-bounds intervals, shared contigs only, merging, removing duplicates, a universe Fisher test, a seeded shuffle) and one complete Python script: A_PATH = "a.bed", B_PATH = "b.bed", gtars RegionSet, an EXPECTED dict of every browser value checked with math.isclose, then the fixes.decision: the text of an earlier interpret run (optional)

Both lanes return the same envelope: lane, verdict, headline, tldr, the lane body, next_steps and prescan_responses. Worked examples: interpret, script, two more readings.

Input fields

Every field is a string.

fieldrequiredmeaning
taskyesinterpret or script.
factsyesA JSON-encoded string with the browser's analysis - see below. Build it with OvKit.buildInput.
titlenoA label for the analysis, up to 160 characters.
contextnoYour notes: what each set is, the assay, the assembly, what you want to conclude. Up to 3,000 characters.
decisionscript onlyPlain text of an earlier interpret run (the page builds it with Recon.decisionText). Up to 6,000 characters.
questionnoAnswered in tldr as a bullet starting "Answer:". Up to 1,200 characters.
retry_notenoOnly on a retry after a malformed reply.

The facts string

settings (the two labels, genome, whether a universe was given, shuffles, seed, and the coordinate, overlap and merge rules); sets (a, b, universe: intervals, invalid lines, contigs and naming style, total and merged bp, widths, self-overlapping intervals, duplicates, unknown contigs, intervals past a chromosome end); pairwise (intersection and union bp, Jaccard, both coverages, overlap coefficient, and for each direction the intervals with an overlap, the nearest-gap bins and the median gap); chance (observed and expected bp, fold, per chromosome, and the shuffle summary with its p-values and result) or null; universe_test (the 2x2 table, odds ratio, p-values, intervals outside the universe, result) or null; metrics (M1..M7); tests (shuffle, universe_fisher); flags (F1.. with severity, category, message and refs); browser_verdict; expected (the values a gtars reproduction must match); and clipped. The intervals themselves are never sent.

Building the body

The simplest way to get a body that matches the page byte for byte is to run the page's own module in Node. overlapkit.js has no dependencies and exports itself with module.exports.

// make-body.js - build the exact body the page sends, with the page's own code.
// Save https://overlap-desk.skillsafe.ai/overlapkit.js next to this file, then:
//   node make-body.js peaks.bed promoters.bed hg38 interpret "TF peaks against promoters" "notes" [universe.bed] > body.json
const fs = require("fs");
const K = require("./overlapkit.js");
const [a, b, genome = "hg38", lane = "interpret", title = "", context = "", universe = ""] = process.argv.slice(2);
const set = { lane, title, context, question: "", decision: "", genome, shuffles: "1000", seed: "42",
  label_a: a.replace(/\.[^.]+$/, ""), label_b: b.replace(/\.[^.]+$/, ""),
  bed_a: fs.readFileSync(a, "utf8"), bed_b: fs.readFileSync(b, "utf8"), universe: universe ? fs.readFileSync(universe, "utf8") : "" };
const X = K.analyze(set);
if (X.empty) throw new Error(X.errors.join("; "));
const body = K.mustBeObject(K.buildInput(X, set));
console.error("browser verdict:", X.hint, "| jaccard:", X.pw.jaccard, "| flags:", X.flags.map(f => f.id + " " + f.category).join(", "));
console.error("idempotency key: overlap-desk:" + lane + ":" + K.hashInput(body) + ":a1");
process.stdout.write(JSON.stringify(body));

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"overlap-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….

The input object IS the request body. There is no {"input": …} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your text.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="overlap-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://overlap-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"overlap-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://overlap-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"overlap-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra and markup_bps 1000 (a 10% markup). hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The body is the input object itself, with no {"input": …} wrapper. /estimate does not validate the body, so check the shape yourself: an object whose every value is a string, task equal to interpret or script, facts non-empty, and facts a JSON string that parses to an object (this is what the page's own guard, OvKit.mustBeObject, refuses to spend without).

# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.js above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task") in ("interpret","script") and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("facts",)) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
#   "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, overlap-desk:<lane>:<hash>:a<attempt> (for example overlap-desk:interpret:yvznap1jd53cr:a1), so a retried request returns the same job instead of billing a second run. Use one key per distinct input: changed sets, background, labels or notes (so changed facts) or a changed reading are a new hash, the same sets in the other lane are a new key, and replaying an old key with a different body is a 409. The page uses OvKit.hashInput(body) for the hash (it covers task, title, context, facts, decision and question; make-body.js prints the key); any stable digest of the body works from other languages. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # interpret or script
KEY="overlap-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"interpret\",\"verdict\":\"fixable\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"lane\":\"interpret\",\"verdict\":\"fixable\",\"headline\":\"The"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is a JSON object serialised as a string. Parse it, then check the lane.

# The reply is a JSON string inside data.output.output. Pull it out and parse it:
printf '%s' "$JOB" | python3 -c 'import sys,json;r=json.loads(json.load(sys.stdin)["output"]["output"]);print(r["verdict"],r["headline"])'

Invariants worth asserting

The output contract

{
  "lane": "interpret" | "script",
  "verdict": "sound" | "caveated" | "unreliable",
  "headline": "...",
  "tldr": ["..."],
  // interpret:
  "metrics": [{"id": "M1", "reading"}],
  "tests": [{"test": "shuffle|universe_fisher", "result": "enriched|depleted|not_significant|not_run", "meaning"}],
  "claims": [{"claim", "support": "supported|partly|not_supported", "why"}],
  "cautions": ["..."],
  // script:
  "fixes": [{"fix", "why", "refs": "F1"}],
  "script": "from gtars.models import RegionSet ...",
  "assumptions": ["..."],
  "checks": ["..."],
  // both:
  "next_steps": ["..."],
  "prescan_responses": [{"ref": "F1", "verdict": "confirmed|dismissed", "note": "..."}]
}

Worked example: interpret

The page's "TF peaks at promoters" example: 250 illustrative transcription-factor peaks against 318 promoter windows on hg38, with a universe of 900 accessible sites. The browser finds both chance tests enriched and raises two flags (promoters that overlap each other, promoters outside the universe), so its read is caveated. The body (facts abbreviated):

{
 "task": "interpret",
 "title": "TF peaks against promoters",
 "context": "Peaks are a transcription factor's ChIP-seq peaks after IDR thresholding. Promoters are TSS +/- 1 kb windows of protein-coding genes. The universe is every accessible site called in the same cells. All on hg38. I want to say this factor binds mostly at promoters.",
 "question": "Is the factor enriched at promoters beyond what accessibility alone explains?",
 "facts": "JSON string of: {\n \"settings\": {\n  \"label_a\": \"TF peaks\",\n  \"label_b\": \"promoters\",\n  \"genome\": \"hg38\",\n  \"genome_label\": \"GRCh38/hg38 primary chromosomes (UCSC chrom.sizes)\",\n  \"universe_given\": true,\n  \"shuffles_requested\": 1000,\n  \"seed\": 42,\n  \"coordinates\": \"BED, 0-based half-open\",\n  \"overlap_rule\": \"two intervals overlap when they share at least one base (book-ended intervals do not)\",\n  \"merge_rule\": \"overlapping and book-ended intervals merge before base-pair metrics (gtars reduce)\"\n },\n \"sets\": {\n  \"a\": {\n   \"label\": \"TF peaks\",\n   \"intervals\": 250,\n   \"merged_bp\": 85947,\n   \"naming_style\": \"ucsc\",\n   \"width_median\": 338,\n   \"intervals_overlapping_own_set\": 2,\n   \"...\": \"more\"\n  },\n  \"b\": {\n   \"label\": \"promoters\",\n   \"intervals\": 318,\n   \"merged_bp\": 616565,\n   \"naming_style\": \"ucsc\",\n   \"width_median\": 2000,\n   \"intervals_overlapping_own_set\": 37,\n   \"...\": \"more\"\n  },\n  \"universe\": {\n   \"label\": \"universe\",\n   \"intervals\": 900,\n   \"merged_bp\": 645065,\n   \"naming_style\": \"ucsc\",\n   \"width_median\": 710,\n   \"intervals_overlapping_own_set\": 10,\n   \"...\": \"more\"\n  }\n },\n \"pairwise\": {\n  \"shared_contigs\": 8,\n  \"shared_contig_names\": [\n   \"chr1\",\n   \"chr2\",\n   \"chr3\",\n   \"chr4\",\n   \"chr5\",\n   \"chr6\",\n   \"chr7\",\n   \"chr8\"\n  ],\n  \"contigs_only_in_a\": [],\n  \"contigs_only_in_b\": [],\n  \"intersection_bp\": 36573,\n  \"union_bp\": 665939,\n  \"jaccard\": 0.0549194,\n  \"coverage_a_by_b\": 0.42553,\n  \"coverage_b_by_a\": 0.0593173,\n  \"overlap_coefficient\": 0.42553,\n  \"a_against_b\": \"{...}\",\n  \"b_against_a\": \"{...}\"\n },\n \"chance\": {\n  \"genome_bp\": 3088286401,\n  \"intervals_tested_a\": 250,\n  \"intervals_tested_b\": 318,\n  \"observed_bp\": 36573,\n  \"expected_bp\": 34.5099,\n  \"fold\": 1060,\n  \"per_contig\": [\n   {\n    \"contig\": \"chr1\",\n    \"n_a\": 43,\n    \"n_b\": 56,\n    \"overlap_bp\": 6429,\n    \"expected_bp\": 5.997,\n    \"fold\": 1072\n   },\n   {\n    \"contig\": \"chr2\",\n    \"n_a\": 45,\n    \"n_b\": 43,\n    \"overlap_bp\": 5908,\n    \"expected_bp\": 5.303,\n    \"fold\": 1114\n   },\n   \"... 6 more\"\n  ],\n  \"shuffle\": {\n   \"shuffles\": 1000,\n   \"seed\": 42,\n   \"null_mean_bp\": 38.624,\n   \"null_sd_bp\": 116.67,\n   \"null_min_bp\": 0,\n   \"null_max_bp\": 787,\n   \"fold\": 946.9,\n   \"z\": 313.1,\n   \"p_greater\": 0.000999,\n   \"p_less\": 1,\n   \"smallest_possible_p\": 0.000999,\n   \"result\": \"enriched\"\n  }\n },\n \"universe_test\": {\n  \"regions\": 900,\n  \"both\": 112,\n  \"only_a\": 141,\n  \"only_b\": 89,\n  \"neither\": 558,\n  \"odds_ratio\": 4.98,\n  \"p_greater\": 2.19e-21,\n  \"p_less\": 1,\n  \"p_two_sided\": 2.385e-21,\n  \"a_outside_universe\": 0,\n  \"b_outside_universe\": 113,\n  \"result\": \"enriched\"\n },\n \"metrics\": [\n  {\n   \"metric\": \"jaccard\",\n   \"value\": 0.0549194,\n   \"basis\": \"shared bases 36573 / bases in either set 665939, both sets merged first\",\n   \"id\": \"M1\"\n  },\n  {\n   \"metric\": \"coverage_a_by_b\",\n   \"value\": 0.42553,\n   \"basis\": \"share of TF peaks's 85947 merged bases that promoters covers\",\n   \"id\": \"M2\"\n  },\n  {\n   \"metric\": \"coverage_b_by_a\",\n   \"value\": 0.0593173,\n   \"basis\": \"share of promoters's 616565 merged bases that TF peaks covers\",\n   \"id\": \"M3\"\n  },\n  {\n   \"metric\": \"overlap_coefficient\",\n   \"value\": 0.42553,\n   \"basis\": \"shared bases / the smaller set's merged bases (85947)\",\n   \"id\": \"M4\"\n  },\n  {\n   \"metric\": \"a_with_overlap\",\n   \"value\": 110,\n   \"fraction\": 0.44,\n   \"basis\": \"110 of 250 intervals of TF peaks share at least one base with promoters\",\n   \"id\": \"M5\"\n  },\n  {\n   \"metric\": \"b_with_overlap\",\n   \"value\": 116,\n   \"fraction\": 0.36478,\n   \"basis\": \"116 of 318 intervals of promoters share at least one base with TF peaks\",\n   \"id\": \"M6\"\n  },\n  {\n   \"metric\": \"a_nearest_b\",\n   \"value\": 1655341.5,\n   \"basis\": \"140 intervals of TF peaks have no overlap; 0 of them lie within 1 kb of promoters, 0 more within 10 kb, and 0 are on contigs promoters never uses; value is their median gap in bp\",\n   \"id\": \"M7\"\n  }\n ],\n \"tests\": [\n  {\n   \"test\": \"shuffle\",\n   \"result\": \"enriched\",\n   \"basis\": \"observed 36573 shared bases against a mean of 38.624 over 1000 within-chromosome shuffles of promoters (fold 946.9, p greater 0.000999, p less 1)\"\n  },\n  {\n   \"test\": \"universe_fisher\",\n   \"result\": \"enriched\",\n   \"basis\": \"universe regions in both 112, only TF peaks 141, only promoters 89, neither 558 (odds ratio 4.98, one-sided p 2.19e-21)\"\n  }\n ],\n \"flags\": [\n  {\n   \"id\": \"F1\",\n   \"severity\": \"medium\",\n   \"category\": \"self_overlap\",\n   \"message\": \"37 of 318 intervals in promoters (B) overlap another interval of the same set; per-interval counts count them separately, while the base-pair metrics use the merged set (299 merged intervals).\",\n   \"refs\": [\n    \"B\"\n   ]\n  },\n  {\n   \"id\": \"F2\",\n   \"severity\": \"medium\",\n   \"category\": \"universe_coverage\",\n   \"message\": \"113 of 318 intervals in promoters overlap no universe region; a universe must contain every region it tests, so the Fisher test ignores them and its background is too narrow.\",\n   \"refs\": [\n    \"B\",\n    \"U\"\n   ]\n  }\n ],\n \"browser_verdict\": \"caveated\",\n \"expected\": {\n  \"n_a\": 250,\n  \"n_b\": 318,\n  \"merged_bp_a\": 85947,\n  \"merged_bp_b\": 616565,\n  \"jaccard\": 0.05491944457,\n  \"coverage_a_by_b\": 0.4255296869,\n  \"coverage_b_by_a\": 0.05931734691,\n  \"overlap_coefficient\": 0.4255296869,\n  \"a_with_overlap\": 110,\n  \"b_with_overlap\": 116,\n  \"n_universe\": 900,\n  \"universe_both\": 112,\n  \"universe_a_only\": 141,\n  \"universe_b_only\": 89,\n  \"universe_neither\": 558,\n  \"fisher_p_greater\": 2.190462047e-21,\n  \"fisher_odds_ratio\": 4.980157781\n },\n \"expected_count\": 17,\n \"clipped\": []\n}"
}

The saved reply (a model following the app's prompt against this exact body; the page replays it for free):

{
 "lane": "interpret",
 "verdict": "caveated",
 "headline": "TF peaks and promoters share far more bases than chance predicts by both the shuffle and universe Fisher tests, but the overlap covers under half of TF peaks's sequence, and self-overlapping promoters plus a universe missing 113 promoters both caveat the numbers.",
 "tldr": [
  "Answer: among the universe's accessible sites, TF peaks and promoters co-occur far above chance (odds ratio 4.98, one-sided p 2.19e-21), but 113 of promoters's 318 intervals sit outside the universe (F2), so that test does not cover all promoters.",
  "Both chance tests call the overlap enriched: 36,573 observed shared bases against a shuffle mean of 38.624 bp over 1,000 within-chromosome shuffles (fold 946.9, p greater 0.000999), plus the universe Fisher result above.",
  "Only 42.6% of TF peaks's 85,947 merged bp and 44% of its 250 intervals (110) overlap promoters, so most TF-peak sequence and most individual peaks sit outside promoters despite the enrichment.",
  "37 of promoters's 318 intervals self-overlap (F1) and 113 lie outside the universe (F2); both shape how the per-interval and universe-test figures should be read."
 ],
 "metrics": [
  {
   "id": "M1",
   "reading": "Jaccard between TF peaks and promoters is 0.0549 (5.49%): the 36,573 bp shared is a small share of the 665,939 bp spanned by the union of the two merged sets, so the two footprints overlap only modestly relative to their combined span."
  },
  {
   "id": "M2",
   "reading": "Coverage_a_by_b is 0.42553 (42.6%): promoters cover 42.6% of TF peaks's 85,947 merged bp (36,573 bp of it), so a substantial minority, not most, of TF-peak sequence lies inside a promoters window."
  },
  {
   "id": "M3",
   "reading": "Coverage_b_by_a is 0.0593173 (5.93%): TF peaks cover only 5.93% of promoters's 616,565 merged bp (36,573 bp of it), so the great majority of promoters's sequence is untouched by TF peaks."
  },
  {
   "id": "M4",
   "reading": "Overlap coefficient is 0.42553 (42.6%): shared bases (36,573 bp) divided by the smaller merged set's bases, TF peaks's 85,947 bp; it equals coverage_a_by_b here because TF peaks is the smaller of the two merged sets."
  },
  {
   "id": "M5",
   "reading": "110 of TF peaks's 250 intervals (44%) share at least one base with a promoters interval; the remaining 140 TF peaks intervals have no promoters overlap at all."
  },
  {
   "id": "M6",
   "reading": "116 of promoters's 318 intervals (36.5%) share at least one base with a TF peaks interval; the remaining 202 promoters intervals have no TF peaks overlap."
  },
  {
   "id": "M7",
   "reading": "Among the 140 TF peaks intervals with no promoters overlap, none sit within 1 kb or 10 kb of a promoters interval and none are on contigs promoters lacks; their median gap to the nearest promoters interval is 1,655,341.5 bp, so these are far misses, not near misses."
  }
 ],
 "tests": [
  {
   "test": "shuffle",
   "result": "enriched",
   "meaning": "Observed 36,573 shared bases far exceed the null mean of 38.624 bp from 1,000 within-chromosome shuffles of promoters (fold 946.9, p greater 0.000999, p less 1), ruling out chance placement within the same chromosomes. This within-chromosome shuffle does not model mappability or accessibility, so on its own it cannot rule out that both sets simply favor generically accessible chromatin."
  },
  {
   "test": "universe_fisher",
   "result": "enriched",
   "meaning": "Among the universe's 900 accessible regions, 112 fall in both TF peaks and promoters versus 141 only in TF peaks, 89 only in promoters and 558 in neither, giving an odds ratio of 4.98 (one-sided p 2.19e-21): co-occurrence at accessible sites is far above chance. This test cannot speak for the 113 promoters intervals outside the universe (F2), and like the shuffle test it shows co-location at accessible sites, not that the factor occupies these promoters functionally."
  }
 ],
 "claims": [
  {
   "claim": "This factor binds mostly at promoters.",
   "support": "not_supported",
   "why": "Only 42.6% of TF peaks's merged bases (M2) and 44% of its 250 intervals, 110 (M5), overlap promoters; both are under half, so most of this factor's peaks and peak sequence sit outside promoters even though the overlap is far above chance."
  },
  {
   "claim": "The factor is enriched at promoters beyond what accessibility alone explains.",
   "support": "partly",
   "why": "The universe Fisher test, run over accessible sites only, shows strong co-occurrence (odds ratio 4.98, one-sided p 2.19e-21), supporting enrichment beyond generic accessibility for sites the universe covers. But 113 of promoters's 318 intervals lie outside the universe (F2), so the test does not speak for all promoters, and the shuffle test used to rule out chance does not itself control for accessibility."
  }
 ],
 "cautions": [
  "Overlap shows TF peaks and promoters share genomic positions; it does not show the factor physically binds, regulates or functions at these promoters beyond what the ChIP-seq peaks already represent.",
  "Enriched in the shuffle and universe tests means more shared bases or co-occurring regions than each test's own background predicts, not that most of the factor's peaks or promoters's sequence overlap.",
  "113 of promoters's 318 intervals are outside the universe and untested by the Fisher test, so the universe result cannot describe those promoters.",
  "These numbers reflect the IDR-thresholded peak calls and the TSS +/- 1 kb promoters windows as pasted; different thresholds or promoters definitions would change every count and fraction here."
 ],
 "next_steps": [
  "Rebuild the universe so all 318 promoters intervals are included (113 currently fall outside it, F2) before re-running the Fisher test, to avoid a narrowed background.",
  "Recompute per-interval promoters statistics on the merged 299-interval promoters set rather than the raw 318, since 37 intervals self-overlap (F1).",
  "Report M2 (42.6% of TF peaks's bases) and M5 (44% of TF peaks's intervals) alongside the enrichment tests so enriched is not read as most binding is at promoters.",
  "When the claim is specifically about accessibility, cite the universe Fisher result (odds ratio 4.98) rather than the within-chromosome shuffle, since only the universe test uses an accessibility-matched background.",
  "Use the script task to generate a reproducible gtars script that applies the F1 and F2 fixes before drawing further conclusions."
 ],
 "prescan_responses": [
  {
   "ref": "F1",
   "verdict": "confirmed",
   "note": "Facts confirm 37 of promoters's 318 intervals self-overlap; base-pair metrics (M2, M3, M4) already use the merged 299-interval set, but per-interval figures (M5, M6) still count the original 318, so this caveat holds as flagged."
  },
  {
   "ref": "F2",
   "verdict": "confirmed",
   "note": "Facts confirm 113 of promoters's 318 intervals fall outside the universe while a_outside_universe is 0, narrowing the Fisher test's background to a subset of promoters, as flagged."
  }
 ]
}

Worked example: script

The same sets with the reading above handed over as decision. The body (facts and decision abbreviated):

{
 "task": "script",
 "title": "TF peaks against promoters",
 "context": "Peaks are a transcription factor's ChIP-seq peaks after IDR thresholding. Promoters are TSS +/- 1 kb windows of protein-coding genes. The universe is every accessible site called in the same cells. All on hg38. I want to say this factor binds mostly at promoters.",
 "question": "",
 "decision": "Verdict: caveated.\nTF peaks and promoters share far more bases than chance predicts by both the shuffle and universe Fisher tests, but the overlap covers under half of TF peaks's sequence, and self-overlapping promoters plus a universe missing 113 promoters both caveat the numbers.\n- M1: Jaccard bet ...",
 "facts": "JSON string of: {\n \"settings\": {\n  \"label_a\": \"TF peaks\",\n  \"label_b\": \"promoters\",\n  \"genome\": \"hg38\",\n  \"genome_label\": \"GRCh38/hg38 primary chromosomes (UCSC chrom.sizes)\",\n  \"universe_given\": true,\n  \"shuffles_requested\": 1000,\n  \"seed\": 42,\n  \"coordinates\": \"BED, 0-based half-open\",\n  \"overlap_rule\": \"two intervals overlap when they share at least one base (book-ended intervals do not)\",\n  \"merge_rule\": \"overlapping and book-ended intervals merge before base-pair metrics (gtars reduce)\"\n },\n \"sets\": {\n  \"a\": {\n   \"label\": \"TF peaks\",\n   \"intervals\": 250,\n   \"merged_bp\": 85947,\n   \"naming_style\": \"ucsc\",\n   \"width_median\": 338,\n   \"intervals_overlapping_own_set\": 2,\n   \"...\": \"more\"\n  },\n  \"b\": {\n   \"label\": \"promoters\",\n   \"intervals\": 318,\n   \"merged_bp\": 616565,\n   \"naming_style\": \"ucsc\",\n   \"width_median\": 2000,\n   \"intervals_overlapping_own_set\": 37,\n   \"...\": \"more\"\n  },\n  \"universe\": {\n   \"label\": \"universe\",\n   \"intervals\": 900,\n   \"merged_bp\": 645065,\n   \"naming_style\": \"ucsc\",\n   \"width_median\": 710,\n   \"intervals_overlapping_own_set\": 10,\n   \"...\": \"more\"\n  }\n },\n \"pairwise\": {\n  \"shared_contigs\": 8,\n  \"shared_contig_names\": [\n   \"chr1\",\n   \"chr2\",\n   \"chr3\",\n   \"chr4\",\n   \"chr5\",\n   \"chr6\",\n   \"chr7\",\n   \"chr8\"\n  ],\n  \"contigs_only_in_a\": [],\n  \"contigs_only_in_b\": [],\n  \"intersection_bp\": 36573,\n  \"union_bp\": 665939,\n  \"jaccard\": 0.0549194,\n  \"coverage_a_by_b\": 0.42553,\n  \"coverage_b_by_a\": 0.0593173,\n  \"overlap_coefficient\": 0.42553,\n  \"a_against_b\": \"{...}\",\n  \"b_against_a\": \"{...}\"\n },\n \"chance\": {\n  \"genome_bp\": 3088286401,\n  \"intervals_tested_a\": 250,\n  \"intervals_tested_b\": 318,\n  \"observed_bp\": 36573,\n  \"expected_bp\": 34.5099,\n  \"fold\": 1060,\n  \"per_contig\": [\n   {\n    \"contig\": \"chr1\",\n    \"n_a\": 43,\n    \"n_b\": 56,\n    \"overlap_bp\": 6429,\n    \"expected_bp\": 5.997,\n    \"fold\": 1072\n   },\n   {\n    \"contig\": \"chr2\",\n    \"n_a\": 45,\n    \"n_b\": 43,\n    \"overlap_bp\": 5908,\n    \"expected_bp\": 5.303,\n    \"fold\": 1114\n   },\n   \"... 6 more\"\n  ],\n  \"shuffle\": {\n   \"shuffles\": 1000,\n   \"seed\": 42,\n   \"null_mean_bp\": 38.624,\n   \"null_sd_bp\": 116.67,\n   \"null_min_bp\": 0,\n   \"null_max_bp\": 787,\n   \"fold\": 946.9,\n   \"z\": 313.1,\n   \"p_greater\": 0.000999,\n   \"p_less\": 1,\n   \"smallest_possible_p\": 0.000999,\n   \"result\": \"enriched\"\n  }\n },\n \"universe_test\": {\n  \"regions\": 900,\n  \"both\": 112,\n  \"only_a\": 141,\n  \"only_b\": 89,\n  \"neither\": 558,\n  \"odds_ratio\": 4.98,\n  \"p_greater\": 2.19e-21,\n  \"p_less\": 1,\n  \"p_two_sided\": 2.385e-21,\n  \"a_outside_universe\": 0,\n  \"b_outside_universe\": 113,\n  \"result\": \"enriched\"\n },\n \"metrics\": [\n  {\n   \"metric\": \"jaccard\",\n   \"value\": 0.0549194,\n   \"basis\": \"shared bases 36573 / bases in either set 665939, both sets merged first\",\n   \"id\": \"M1\"\n  },\n  {\n   \"metric\": \"coverage_a_by_b\",\n   \"value\": 0.42553,\n   \"basis\": \"share of TF peaks's 85947 merged bases that promoters covers\",\n   \"id\": \"M2\"\n  },\n  {\n   \"metric\": \"coverage_b_by_a\",\n   \"value\": 0.0593173,\n   \"basis\": \"share of promoters's 616565 merged bases that TF peaks covers\",\n   \"id\": \"M3\"\n  },\n  {\n   \"metric\": \"overlap_coefficient\",\n   \"value\": 0.42553,\n   \"basis\": \"shared bases / the smaller set's merged bases (85947)\",\n   \"id\": \"M4\"\n  },\n  {\n   \"metric\": \"a_with_overlap\",\n   \"value\": 110,\n   \"fraction\": 0.44,\n   \"basis\": \"110 of 250 intervals of TF peaks share at least one base with promoters\",\n   \"id\": \"M5\"\n  },\n  {\n   \"metric\": \"b_with_overlap\",\n   \"value\": 116,\n   \"fraction\": 0.36478,\n   \"basis\": \"116 of 318 intervals of promoters share at least one base with TF peaks\",\n   \"id\": \"M6\"\n  },\n  {\n   \"metric\": \"a_nearest_b\",\n   \"value\": 1655341.5,\n   \"basis\": \"140 intervals of TF peaks have no overlap; 0 of them lie within 1 kb of promoters, 0 more within 10 kb, and 0 are on contigs promoters never uses; value is their median gap in bp\",\n   \"id\": \"M7\"\n  }\n ],\n \"tests\": [\n  {\n   \"test\": \"shuffle\",\n   \"result\": \"enriched\",\n   \"basis\": \"observed 36573 shared bases against a mean of 38.624 over 1000 within-chromosome shuffles of promoters (fold 946.9, p greater 0.000999, p less 1)\"\n  },\n  {\n   \"test\": \"universe_fisher\",\n   \"result\": \"enriched\",\n   \"basis\": \"universe regions in both 112, only TF peaks 141, only promoters 89, neither 558 (odds ratio 4.98, one-sided p 2.19e-21)\"\n  }\n ],\n \"flags\": [\n  {\n   \"id\": \"F1\",\n   \"severity\": \"medium\",\n   \"category\": \"self_overlap\",\n   \"message\": \"37 of 318 intervals in promoters (B) overlap another interval of the same set; per-interval counts count them separately, while the base-pair metrics use the merged set (299 merged intervals).\",\n   \"refs\": [\n    \"B\"\n   ]\n  },\n  {\n   \"id\": \"F2\",\n   \"severity\": \"medium\",\n   \"category\": \"universe_coverage\",\n   \"message\": \"113 of 318 intervals in promoters overlap no universe region; a universe must contain every region it tests, so the Fisher test ignores them and its background is too narrow.\",\n   \"refs\": [\n    \"B\",\n    \"U\"\n   ]\n  }\n ],\n \"browser_verdict\": \"caveated\",\n \"expected\": {\n  \"n_a\": 250,\n  \"n_b\": 318,\n  \"merged_bp_a\": 85947,\n  \"merged_bp_b\": 616565,\n  \"jaccard\": 0.05491944457,\n  \"coverage_a_by_b\": 0.4255296869,\n  \"coverage_b_by_a\": 0.05931734691,\n  \"overlap_coefficient\": 0.4255296869,\n  \"a_with_overlap\": 110,\n  \"b_with_overlap\": 116,\n  \"n_universe\": 900,\n  \"universe_both\": 112,\n  \"universe_a_only\": 141,\n  \"universe_b_only\": 89,\n  \"universe_neither\": 558,\n  \"fisher_p_greater\": 2.190462047e-21,\n  \"fisher_odds_ratio\": 4.980157781\n },\n \"expected_count\": 17,\n \"clipped\": []\n}"
}

The saved reply. Its script was run with gtars 0.9.2 and scipy 1.15.3 on the example's a.bed, b.bed and universe.bed: the reproduction check matched all 17 expected values and the seeded shuffle ran.

{
 "lane": "script",
 "verdict": "caveated",
 "headline": "TF peaks and promoters share far more bases than chance predicts by both tests, but coverage is partial and two flags (self-overlapping promoters, a narrowed universe) still caveat the numbers this script reproduces.",
 "tldr": [
  "Verdict caveated: the shuffle and universe Fisher tests both call TF peaks/promoters enriched, but coverage_a_by_b (M2, 42.6%) and a_with_overlap (M5, 44%) show partial coverage, not majority binding.",
  "The script reproduces gtars's own numbers (jaccard, coverage, overlap coefficient, both overlap counts, the universe Fisher test) and checks each against facts.expected before applying fixes.",
  "Fix F1 merges promoters with reduce() so the 37 self-overlapping intervals do not get read as the same denominator as the 299 merged intervals used for base-pair metrics.",
  "Fix F2 flags that 113 of promoters's 318 intervals sit outside the given universe, so the Fisher test's odds ratio (4.98) cannot speak for them until a wider universe is supplied.",
  "The seeded shuffle test (seed 42, 1000 shuffles) is reproduced as a direction check only; it is not expected to match the browser's exact p-value."
 ],
 "fixes": [
  {
   "fix": "Merge promoters with RegionSet.reduce() before reporting per-interval counts, so the 37 self-overlapping intervals (F1) are folded into the 299 merged intervals the base-pair numbers already use.",
   "why": "Raw per-interval counts (318) and merged base-pair counts (299 merged intervals) use different denominators; without merging first, self-overlaps make the two look inconsistent.",
   "refs": "F1"
  },
  {
   "fix": "Keep the Fisher test on the given 900-region universe but print that it is silent on the 113 promoters intervals outside it (F2), leaving AUTHOR_INPUT_NEEDED for a wider universe.bed before rerunning.",
   "why": "A universe test can only speak for regions inside the universe; no amount of script logic can include the 113 missing intervals without a new file from the user.",
   "refs": "F2"
  },
  {
   "fix": "Reproduce the seeded within-chromosome shuffle (random.seed(42), 1000 shuffles) as a direction check on the observed shared bases, not a claim of matching the browser's exact p-value.",
   "why": "The house rules require a seeded shuffle reproduction step and forbid claiming to reproduce the browser's stochastic p-value exactly.",
   "refs": ""
  }
 ],
 "script": "import math\nimport random\nimport statistics\nimport collections\nfrom pathlib import Path\nimport scipy.stats\nfrom gtars.models import RegionSet\n\nA_PATH = \"a.bed\"\nB_PATH = \"b.bed\"\nU_PATH = \"universe.bed\"\n\nfor path in (A_PATH, B_PATH, U_PATH):\n    if not Path(path).is_file():\n        raise FileNotFoundError(f\"missing input file: {path}\")\n\na = RegionSet(A_PATH)  # TF peaks\nb = RegionSet(B_PATH)  # promoters\nu = RegionSet(U_PATH)  # universe\n\nEXPECTED = {\n    \"n_a\": 250,\n    \"n_b\": 318,\n    \"merged_bp_a\": 85947,\n    \"merged_bp_b\": 616565,\n    \"jaccard\": 0.05491944457,\n    \"coverage_a_by_b\": 0.4255296869,\n    \"coverage_b_by_a\": 0.05931734691,\n    \"overlap_coefficient\": 0.4255296869,\n    \"a_with_overlap\": 110,\n    \"b_with_overlap\": 116,\n    \"n_universe\": 900,\n    \"universe_both\": 112,\n    \"universe_a_only\": 141,\n    \"universe_b_only\": 89,\n    \"universe_neither\": 558,\n    \"fisher_p_greater\": 2.190462047e-21,\n    \"fisher_odds_ratio\": 4.980157781,\n}\n\na_merged = a.reduce()\nb_merged = b.reduce()\na_hits_b = a.any_overlaps(b)\nb_hits_a = b.any_overlaps(a)\nu_hits_a = u.any_overlaps(a)\nu_hits_b = u.any_overlaps(b)\n\nboth = sum(1 for ha, hb in zip(u_hits_a, u_hits_b) if ha and hb)\nonly_a = sum(1 for ha, hb in zip(u_hits_a, u_hits_b) if ha and not hb)\nonly_b = sum(1 for ha, hb in zip(u_hits_a, u_hits_b) if hb and not ha)\nneither = sum(1 for ha, hb in zip(u_hits_a, u_hits_b) if not ha and not hb)\ntable = [[both, only_a], [only_b, neither]]\nodds_ratio, p_greater = scipy.stats.fisher_exact(table, alternative=\"greater\")\n\nOBSERVED = {\n    \"n_a\": len(a),\n    \"n_b\": len(b),\n    \"merged_bp_a\": a_merged.get_nucleotide_length(),\n    \"merged_bp_b\": b_merged.get_nucleotide_length(),\n    \"jaccard\": a.jaccard(b),\n    \"coverage_a_by_b\": a.coverage(b),\n    \"coverage_b_by_a\": b.coverage(a),\n    \"overlap_coefficient\": a.overlap_coefficient(b),\n    \"a_with_overlap\": sum(a_hits_b),\n    \"b_with_overlap\": sum(b_hits_a),\n    \"n_universe\": len(u),\n    \"universe_both\": both,\n    \"universe_a_only\": only_a,\n    \"universe_b_only\": only_b,\n    \"universe_neither\": neither,\n    \"fisher_p_greater\": p_greater,\n    \"fisher_odds_ratio\": odds_ratio,\n}\n\nprint(\"Reproduction check against the browser (facts.expected):\")\nmismatches = 0\nfor key, want in EXPECTED.items():\n    got = OBSERVED[key]\n    if not math.isclose(got, want, rel_tol=1e-6):\n        mismatches += 1\n        print(f\"  MISMATCH {key}: script={got!r} browser={want!r}\")\nprint(f\"  {mismatches} mismatch(es)\" if mismatches else \"  all EXPECTED keys matched\")\n\n# --- Fix F1: 37 of promoters's 318 raw intervals overlap another promoters\n# interval. Per-interval counts (like b_with_overlap above) use the raw 318;\n# base-pair counts already use the merged 299. Report both so they are never\n# read as the same denominator.\nprint(f\"F1 fix: promoters raw intervals={len(b)}, merged intervals={len(b_merged)}, \"\n      f\"self-overlapping intervals folded into the merge=37\")\n\n# --- Fix F2: the given universe leaves 113 of promoters's 318 intervals with\n# no test coverage (facts.universe_test.b_outside_universe). This count is\n# copied from the browser's flag, not recomputed here.\nPROMOTERS_OUTSIDE_UNIVERSE = 113\nprint(f\"F2 fix: the Fisher test above is silent on \"\n      f\"{PROMOTERS_OUTSIDE_UNIVERSE} of promoters's intervals; \"\n      f\"AUTHOR_INPUT_NEEDED: supply a wider universe.bed covering them, \"\n      f\"then rerun scipy.stats.fisher_exact above.\")\n\n# --- Reproduction step: seeded within-chromosome shuffle of promoters (B) ---\n# Chromosome bounds are approximated from the two pasted sets' own coordinates.\n# AUTHOR_INPUT_NEEDED: a real hg38 chrom.sizes path would give true bounds;\n# without one this null is only indicative, not the browser's exact null.\nrandom.seed(42)\nbounds = collections.defaultdict(lambda: [None, None])\nfor region_set in (a, b):\n    for r in region_set:\n        lo, hi = bounds[r.chr]\n        bounds[r.chr] = [r.start if lo is None else min(lo, r.start),\n                          r.end if hi is None else max(hi, r.end)]\n\nb_chrs = [r.chr for r in b]\nb_widths = [r.end - r.start for r in b]\nnull_bp = []\nfor _ in range(1000):\n    s_chrs, s_starts, s_ends = [], [], []\n    for chrom, width in zip(b_chrs, b_widths):\n        lo, hi = bounds[chrom]\n        span = max(hi - lo - width, 1)\n        start = lo + random.randrange(0, span)\n        s_chrs.append(chrom)\n        s_starts.append(start)\n        s_ends.append(start + width)\n    b_shuffled = RegionSet.from_vectors(s_chrs, s_starts, s_ends)\n    null_bp.append(a.coverage(b_shuffled) * OBSERVED[\"merged_bp_a\"])\n\nnull_mean = statistics.mean(null_bp)\nnull_sd = statistics.pstdev(null_bp)\nprint(f\"Shuffle reproduction (seed=42, 1000 shuffles): \"\n      f\"null_mean_bp~={null_mean:.3f}, null_sd_bp~={null_sd:.3f}\")\nprint(\"This direction-checks the browser's enrichment call (observed far above \"\n      \"this null); it does not reproduce the browser's exact null distribution, \"\n      \"seed stream, or p-value, since chromosome bounds here are approximated.\")\n",
 "assumptions": [
  "a.bed, b.bed and universe.bed are the tab-separated BED files the page wrote for TF peaks, promoters and the universe, already 0-based half-open per settings.coordinates.",
  "gtars 0.9.2's RegionSet, reduce(), jaccard(), coverage(), overlap_coefficient() and any_overlaps() behave as the browser's own computation used them.",
  "No chrom.sizes file is available in this environment, so the shuffle test's chromosome bounds are approximated from the pasted sets' own coordinates, not the true hg38 assembly.",
  "The universe file itself is already available (settings.universe_given is true); only a wider replacement for the F2 fix needs AUTHOR_INPUT_NEEDED."
 ],
 "checks": [
  "The reproduction block prints no MISMATCH line for any of the 17 facts.expected keys.",
  "F1's printed line shows merged intervals=299 against raw intervals=318, confirming the 37 self-overlaps were folded in.",
  "F2's printed line still names 113 uncovered promoters intervals until a wider universe.bed is supplied and the Fisher block is rerun.",
  "The shuffle's printed null_mean_bp is far below the observed 36,573 shared bases, matching the browser's enriched call in direction only."
 ],
 "next_steps": [
  "Run the script with a.bed, b.bed and universe.bed in place and confirm the reproduction check prints no mismatch against facts.expected.",
  "Supply a chrom.sizes path (AUTHOR_INPUT_NEEDED) so the shuffle test uses real hg38 chromosome bounds instead of the approximation built from the pasted sets.",
  "Rebuild universe.bed to include the 113 promoters intervals it currently misses (F2), then rerun the Fisher block.",
  "Report coverage_a_by_b (M2, 42.6%) and a_with_overlap (M5, 44%) alongside the enrichment tests so enriched is not read as most binding is at promoters.",
  "Use the merged 299-interval promoters set (b_merged) for any further per-interval analysis so self-overlap (F1) does not double count."
 ],
 "prescan_responses": [
  {
   "ref": "F1",
   "verdict": "confirmed",
   "note": "Confirmed: 37 of promoters's 318 raw intervals overlap another promoters interval; the script merges with reduce() (299 intervals) before any per-interval fix so these are not double counted."
  },
  {
   "ref": "F2",
   "verdict": "confirmed",
   "note": "Confirmed: 113 of promoters's 318 intervals fall outside the given universe, so the Fisher test above is silent on them until a wider universe.bed is supplied (AUTHOR_INPUT_NEEDED)."
  }
 ]
}

Two more readings

"Two replicates": the browser's read is sound (no flags; both tests enriched). The saved reply chose a tighter verdict, which is allowed - it may never be looser:

{
 "lane": "interpret",
 "verdict": "caveated",
 "headline": "rep1 and rep2 share 67668 bp of merged sequence (Jaccard 0.622, coverage 76.5% rep1-by-rep2 and 77.0% rep2-by-rep1, with 85.4% and 86.3% of intervals overlapping) far beyond a shuffle-test chance null of 5 bp, but the universe used for the Fisher enrichment test already contains both replicates in full, so that enrichment figure is not independent evidence.",
 "tldr": [
  "Answer: rep1 and rep2 overlap far beyond chance on both the shuffle test and the universe Fisher test, with coverage about three-quarters in each direction (76.5% rep1-by-rep2, 77.0% rep2-by-rep1) - concordant, but the universe test's background already contains both replicates, so treat concordant enough to merge as a threshold call these numbers alone do not settle.",
  "Jaccard is 0.622 (67668 shared bp of 108746 bp either set covers once merged); 316 of rep1's 370 intervals (85.4%) and 316 of rep2's 366 intervals (86.3%) share at least one base with the other set.",
  "The within-chromosome shuffle test (1000 shuffles, seed 42, null mean 5 bp, fold 13530, p_greater 0.000999) and the universe Fisher test (900 regions, odds ratio 56.18, p 5.557e-127) both call the overlap enriched.",
  "54 of rep1's 370 intervals and 50 of rep2's 366 intervals never overlap the other set, with median gaps of 1594046 bp and 1667915.5 bp - these look like distinct calls, not near misses.",
  "The Fisher test's universe is the merged peak list of every sample in the batch, and both rep1 and rep2 fall entirely inside it (0 regions of either outside), so that background is not independent of the sets being compared."
 ],
 "metrics": [
  {
   "id": "M1",
   "reading": "Jaccard similarity between rep1 and rep2 is 0.622 (62.2%): the 67668 bp the two sets share divided by the 108746 bp either set covers once both are merged."
  },
  {
   "id": "M2",
   "reading": "Coverage of rep1 by rep2 is 0.765 (76.5%): of rep1's 88493 merged bases, that fraction is also covered by rep2."
  },
  {
   "id": "M3",
   "reading": "Coverage of rep2 by rep1 is 0.770 (77.0%): of rep2's 87921 merged bases, that fraction is also covered by rep1."
  },
  {
   "id": "M4",
   "reading": "The overlap coefficient is 0.770 (77.0%): the 67668 shared bases divided by rep2's 87921 merged bases, the smaller of the two sets."
  },
  {
   "id": "M5",
   "reading": "316 of rep1's 370 intervals (85.4%) share at least one base with rep2."
  },
  {
   "id": "M6",
   "reading": "316 of rep2's 366 intervals (86.3%) share at least one base with rep1."
  },
  {
   "id": "M7",
   "reading": "Among rep1's 54 non-overlapping intervals, the median gap to the nearest rep2 interval is 1594046 bp; none of those 54 lie within 1 kb or 10 kb of rep2 and none are on contigs rep2 lacks, so the non-overlapping rep1 intervals sit far from any rep2 peak rather than just missing one nearby."
  }
 ],
 "tests": [
  {
   "test": "shuffle",
   "result": "enriched",
   "meaning": "The 67668 bp rep1 and rep2 share is far above the within-chromosome shuffle null (mean 5 bp, sd 27.8662, over 1000 shuffles, seed 42), fold 13530, p_greater 0.000999, the smallest p this shuffle count can report. This rules out that so much shared sequence arose merely from placing sets of this size on these 8 chromosomes at random, but the shuffle keeps each contig's own length and does not model mappability, chromatin accessibility, or CTCF motif structure, so it cannot rule out that shared accessible regions, not only shared CTCF signal, explain part of the overlap."
  },
  {
   "test": "universe_fisher",
   "result": "enriched",
   "meaning": "Within the 900-region universe, 316 regions carry both rep1 and rep2, 54 only rep1, 50 only rep2, and 480 neither; the one-sided Fisher test gives odds ratio 56.18, p 5.557e-127. Both rep1 and rep2 fall entirely inside this universe (0 regions of either outside it), and the universe itself is the merged peak list of every sample in the batch, so this enrichment is only as independent of rep1 and rep2 as that universe is - it cannot rule out that the universe's own construction from batch peaks inflates the co-occurrence rate."
  }
 ],
 "claims": [
  {
   "claim": "The replicates are concordant.",
   "support": "supported",
   "why": "Jaccard is 0.622, coverage is 76.5% (rep1 by rep2) and 77.0% (rep2 by rep1), 85.4% of rep1's intervals and 86.3% of rep2's intervals overlap the other set, and the within-chromosome shuffle test calls the shared 67668 bp enriched (fold 13530, p_greater 0.000999) against a neutral random background."
  },
  {
   "claim": "The replicates are concordant enough to merge.",
   "support": "partly",
   "why": "Coverage (76.5% and 77.0%) and interval overlap (85.4% and 86.3%) are strong and the shuffle test calls the overlap enriched, but 54 of rep1's 370 intervals and 50 of rep2's 366 intervals show no overlap at all (median gaps 1594046 bp and 1667915.5 bp), and the universe Fisher test's odds ratio 56.18 is measured against a universe that already fully contains both replicates (0 regions of either outside it), so it cannot add independent support; the facts do not state a coverage threshold for enough to merge."
  }
 ],
 "cautions": [
  "Overlap between rep1 and rep2 shows shared genomic positions, not a shared regulatory mechanism, and does not by itself show the same CTCF event in the same cells beyond what the context states.",
  "The universe used for the Fisher test is the merged peak list of every sample in the batch, and both rep1 and rep2 fall entirely inside it (0 regions of either outside), so its enrichment result is not measured against a background independent of the sets being compared.",
  "The within-chromosome shuffle null does not model mappability, chromatin accessibility, or CTCF motif structure, so it cannot rule out that shared accessible or mappable regions, not only shared CTCF calling, explain part of the enrichment.",
  "54 of rep1's 370 intervals and 50 of rep2's 366 intervals never overlap the other set, with median gaps of 1594046 bp and 1667915.5 bp, so these look like distinct calls rather than near-misses."
 ],
 "next_steps": [
  "Treat the shuffle-test enrichment (fold 13530, p_greater 0.000999, a neutral within-chromosome random background) as the primary chance check, and treat the universe Fisher result (odds ratio 56.18) as supporting but not independent, since the universe already fully contains both replicates.",
  "Set an explicit coverage or jaccard threshold for concordant enough to merge before deciding, since jaccard 0.622 and coverage 76.5%/77.0% are strong but not complete.",
  "Inspect the 54 rep1-only and 50 rep2-only non-overlapping intervals (median gaps 1594046 bp and 1667915.5 bp) to see whether they are true replicate-specific calls before merging.",
  "If a merged consensus is produced, base it on gtars reduce() over the overlapping pair, keeping the universe Fisher result as secondary support rather than the primary enrichment claim.",
  "When reporting concordance, cite the shuffle-test numbers (fold 13530, p 0.000999) alongside coverage (76.5%/77.0%), not the universe odds ratio (56.18) alone, given the universe's dependence on the batch's own peaks."
 ],
 "prescan_responses": []
}

"Zero overlap": hg19 enhancers with UCSC names against Ensembl-named windows, genome set to hg38. The browser's read is unreliable and neither test can run:

{
 "lane": "interpret",
 "verdict": "unreliable",
 "headline": "The zero overlap between enhancers and GWAS windows reflects incompatible inputs, not genomic distance, because the two sets use different contig names and likely different genome assemblies.",
 "tldr": [
  "Answer: The overlap is zero because enhancers uses ucsc contig names (chr1, chr2) while GWAS windows uses ensembl names (1, 2), so no contig name matches between the sets, and 10 enhancer intervals also extend past hg38 chromosome ends but fit hg19, pointing to a build mismatch.",
  "0 of 160 enhancer intervals and 0 of 120 GWAS window intervals share any base (M5, M6), and jaccard, both coverage metrics and the overlap coefficient are all 0 (M1-M4).",
  "Neither the shuffle test nor a universe test could be run, so there is no comparison to chance available for this pair (both tests not_run).",
  "6 of 120 GWAS window intervals overlap another interval in the same set, so its 120 per-interval count differs from its merged 117-interval, 2,368,300 bp set used for base-pair metrics.",
  "Verdict is unreliable: fix contig naming and confirm enhancers' genome build before treating this zero as a biological finding."
 ],
 "metrics": [
  {
   "id": "M1",
   "reading": "Jaccard between enhancers (A) and GWAS windows (B) is 0: of the 2,639,879 bp spanned by either set after merging, none is shared, so as computed the sets occupy no common bases."
  },
  {
   "id": "M2",
   "reading": "Coverage of enhancers (A) by GWAS windows (B) is 0: none of enhancers' 271,579 merged bases are covered by GWAS windows."
  },
  {
   "id": "M3",
   "reading": "Coverage of GWAS windows (B) by enhancers (A) is 0: none of GWAS windows' 2,368,300 merged bases are covered by enhancers."
  },
  {
   "id": "M4",
   "reading": "The overlap coefficient is 0, shared bases over the smaller set's 271,579 merged bases (enhancers), confirming no shared bases no matter which set is used as the denominator."
  },
  {
   "id": "M5",
   "reading": "0 of the 160 enhancer (A) intervals share even one base with a GWAS window (B), a 0% interval-level overlap fraction."
  },
  {
   "id": "M6",
   "reading": "0 of the 120 GWAS window (B) intervals share even one base with an enhancer (A), also a 0% interval-level overlap fraction."
  },
  {
   "id": "M7",
   "reading": "The median gap from enhancers to the nearest GWAS window has no value: all 160 non-overlapping enhancer intervals sit on contigs GWAS windows never uses, with 0 within 1 kb and 0 more within 10 kb, so no distance could be measured."
  }
 ],
 "tests": [
  {
   "test": "shuffle",
   "result": "not_run",
   "meaning": "The shuffle test did not run because no interval from either set sits on a contig the genome track lists in a form both sets share; the naming mismatch blocks building a within-chromosome null, so chance cannot be assessed here. Renaming GWAS windows' contigs to match the genome, and confirming enhancers is on hg38, would let it run."
  },
  {
   "test": "universe_fisher",
   "result": "not_run",
   "meaning": "No universe of background regions was supplied, so a Fisher's exact test against a background could not be run; nothing here says whether the pattern differs from a defined background set by chance. It would need a universe BED file with settings.universe_given true."
  }
 ],
 "claims": [
  {
   "claim": "Most GWAS windows should hit an enhancer",
   "support": "not_supported",
   "why": "As computed, 0 of the 120 GWAS windows share any base with an enhancer (M6), but this pair cannot really test that expectation: enhancers' ucsc contig names and GWAS windows' ensembl contig names never match exactly (F3), and enhancers' coordinates fit hg19 rather than the requested hg38 (F1), so the zero reflects broken inputs rather than a biological answer."
  }
 ],
 "cautions": [
  "A zero here does not by itself mean these enhancers and SNP windows are apart in the genome, since incompatible contig names and a possible genome-build mismatch can force intersection_bp to 0 on their own.",
  "Even a nonzero overlap would only show that two regions share genomic positions, not that any enhancer regulates or drives a gene near a GWAS-associated SNP.",
  "No universe test was run, so nothing here says whether any true overlap would exceed what a defined background of accessible or tested regions would predict.",
  "The self-overlapping intervals in GWAS windows (F4) mean interval-level counts and merged base-pair totals describe slightly different versions of that set."
 ],
 "next_steps": [
  "Rename GWAS windows' contigs to the ucsc style (chr1, chr2, ...) that enhancers already uses, or the reverse, so shared chromosomes can be found.",
  "Confirm which assembly enhancers is actually on: its out-of-bounds hg38 coordinates fit GRCh37/hg19 (F1), so lift over to hg38 or re-run the comparison on hg19.",
  "After fixing naming and assembly, re-run the overlap and the within-chromosome shuffle test to see whether shared bases exceed chance.",
  "Merge GWAS windows' 6 self-overlapping intervals (F4) into its 117-interval set before re-reading interval-level overlap counts.",
  "Add a universe of background regions if enrichment over a defined background is the question, so a Fisher's exact test can run."
 ],
 "prescan_responses": [
  {
   "ref": "F1",
   "verdict": "confirmed",
   "note": "10 enhancer intervals extend past hg38 chromosome lengths (e.g. chr1 max_end 249,199,252 vs 248,956,422 bp) but fit GRCh37/hg19, so enhancers is likely on the wrong assembly for this hg38 comparison."
  },
  {
   "ref": "F2",
   "verdict": "confirmed",
   "note": "None of GWAS windows' contigs (ensembl-style 1, 2, 9, 11) are in the hg38 genome's ucsc-style names, so this set cannot be compared to chance at all."
  },
  {
   "ref": "F3",
   "verdict": "confirmed",
   "note": "Enhancers uses ucsc names (chr1, chr2, chr9) and GWAS windows uses ensembl names (1, 2, 9); since contigs are matched exactly, this alone forces intersection_bp to 0 regardless of true positions."
  },
  {
   "ref": "F4",
   "verdict": "confirmed",
   "note": "6 of 120 GWAS window intervals overlap another interval in the same set, so its 120 per-interval count differs from the merged 117-interval, 2,368,300 bp set used for base-pair metrics."
  }
 ]
}

Truncation and partial results

If your balance sits between min_credits and hold_credits, the run still executes with a smaller output cap and the job carries "truncated": true. The JSON may then stop mid-object: close it (the page's Recon.closeJson does this) and show the sections that arrived, saying how many of the lane's sections were recovered, rather than treating a clipped reply as complete.