Monitoring
Method and ARKA's own published figures for certainty burn-down, calibration status, drift under three monitoring strategies, the four appropriateness dimensions, and the corrections feed. Customer-specific rates stay behind an authenticated packet.
Last updated: August 22, 2026
v1.0.0 · 22 August 2026 · as of 2026-08-22
This public report carries no customer names and no site-specific customer rates.
Full report PDF (byte-identical on regenerate): /api/monitoring/report
Certainty burn-down (volume-weighted)
As of 2026-08-22 · matrix 2.0.0 · generated from docs/certainty-coverage.json
The knowledge base is a strong prototype and an inadequate production state — of 1536 matrix ratings, 40 are high, 119 moderate, 22 low, and 1355 very low GRADE certainty, with 21 of 1536 abstaining — and Volume III said advisors forgive a small knowledge base and do not forgive a founder who claims completeness.
Current distribution — raw, nominal volume-weighted, setting-adjusted
| Certainty | Raw | Nominal volume-weighted | Setting-adjusted |
|---|---|---|---|
high | 40 | 1694 | 1694 |
moderate | 119 | 5130 | 5130 |
low | 22 | 1764 | 1764 |
very_low | 1355 | 63636 | 63636 |
The volume-weighted very_low share is 88.1%, below the raw share of 88.2%, because the knowledge base was built from the most common scenarios outward — higher-certainty ratings concentrate where order volume is highest.
Care-setting review has not yet been done for 0 of 1536 ratings covering 0% of order volume. Until it is, those ratings are reported one GRADE level lower than their nominal grade. Item 36 is the work that changes this, and here is its schedule: Session 1 (2026-09-11, 90 min); Session 2 (2026-09-25, 90 min); Session 3 (2026-10-09, 90 min); Session 4 (2026-10-23, 90 min).
Schedule and reviews
- Ratings scheduled for panel review in the next 90 days: 50
- Reviewed to date: 6 (raised 0, lowered 6, unchanged 0)
History (each prior emit)
2026-08-22· matrix2.0.0· raw very_low 1349 · weighted very_low 63534 · commit7427e2fb88c82026-08-22· matrix2.0.0· raw very_low 1349 · weighted very_low 63534 · commit7427e2fb88c82026-08-22· matrix2.0.0· raw very_low 1349 · weighted very_low 65298 · commit7427e2fb88c82026-08-23· matrix2.0.0· raw very_low 1349 · weighted very_low 65298 · commitd3405a09a2e72026-08-23· matrix2.0.0· raw very_low 1355 · weighted very_low 65400 · commitd3405a09a2e72026-08-23· matrix2.0.0· raw very_low 1355 · weighted very_low 63658 · commitd3405a09a2e72026-08-23· matrix2.0.0· raw very_low 1355 · weighted very_low 63636 · commitd3405a09a2e7
No future trend line is projected — four sessions are scheduled and no rate data yet exists. A projection here would be a forecast dressed as a measurement.
Calibration — not yet evaluated
To evaluate calibration in this population, provide at least 100 paired rows with predictedProbability (0–1), observedBinaryOutcome (0 or 1), and careSetting (ED | inpatient | ambulatory | virtual | unknown). Synthetic benchmarks are for reproducibility only and never fill this panel.
- Required fields:
predictedProbability, observedBinaryOutcome, careSetting - Minimum paired outcomes: 100
- Cases supplied without outcomes: 0
Outcome-signal ledger
Claims-derived observed facts for each scored (or claim-proxied) imaging order. Provenance=measured. Predicted probabilities from ARKA never set these bits.
Calibration and drift figures on this page are weighted by arm-exposure covariates (justification, default-order, and peer-packet exposure) so that a change in who is scored after ARKA changes behaviour is not read as calibration drift. Unweighted, adherence-weighted, and sampling-weighted strategies are shown together; the unweighted figure is never published alone once an intervention is live (prompt 34.4).
Downstream utilization within 90 days is a utilization fact read from claims. It is not harm, not an adverse event, and not a quality adjudication.
| Signal | Window | n | Observed rate |
|---|---|---|---|
study_performed | 90d | 115 | 1.000 |
downstream_utilization_90d | 90d | 115 | 0.478 |
Drift — three strategies
Site arka-method · 2026-06 · provenance=illustrative · intervention live · arm exposure 0.400
Population, label, and performance drift under unweighted, adherence-weighted, and sampling-weighted strategies. An unweighted figure is never published alone once an intervention is live.
| Strategy | Population PSI | Label delta | ECE delta | Discrimination |
|---|---|---|---|---|
unweightedNaïve average over observed rows. Biased once an intervention changes who is scored or labelled; published with an explicit caution when interventionLive is true. Intervention is live at this site — the unweighted figure is contaminated by the feedback loop and must be read beside the adherence-weighted and sampling-weighted figures, not alone. | 0.1408 | 0.1450 | 0.0875 | 0.6800 |
adherence_weightedReweights by the inverse of arm-adherence probability so evaluated rows resemble the pre-intervention population under selective uptake. | 0.1408 | 0.1450 | 0.0875 | 0.6800 |
sampling_weightedReweights by the inverse of outcome-sampling probability so label drift is not confused with selective outcome capture. | 0.1408 | 0.1450 | 0.0875 | 0.6800 |
Appropriateness — four dimensions
Not a single scalar. Shared decision and patient understanding stay not_measured until those instruments are fielded. Public surface shows ARKA's certainty-census evidence adequacy; site indication rates remain in authenticated packets.
- Indication supported
- not_measured · illustrative
- Evidence adequate
- not met · 181/1536 (11.8%) · measured
- Shared decision occurred
- not_measured · illustrative
- Patient understanding
- not_measured — three questions, at the end of the encounter, on whatever survey the site already runs
Corrections
We have been wrong in public. This is the list. Adding a row is a code change with a test, not an edit to a page, so the log cannot drift from what we actually retracted.
5 published. The count only goes up. Machine artefact: docs/corrections.json.
We named a training simulator ARKA-ED (historical; the product is ARKA-SIM) on the public site, including an FDA-notice clause that marketed it for time-critical emergency imaging decisions.
- Where
- Public marketing site and lib/compliance/fda-notice-copy.ts — regulatory audit 2026-07-22 Finding F-01; Volume III §25.1 item 13.
- When
- 2026-07-22 (July 22, 2026)
- Why it was wrong
- ARKA-ED (historical; the product is ARKA-SIM) reads as a deployed emergency-department product we do not have. The time-critical clause is the exact trigger FDA's January 2026 CDS guidance uses to pull software back into device regulation. We found this in our own audit, not a regulator's.
- What we say now
- The product is ARKA-SIM, a training simulator. No user-facing string names a live emergency-department product. The FDA-notice string states that the simulator is not used for live, time-critical emergency decisions.
- Commit that fixed it
- cd377bb1854cdde9fa5fcb377e96e01d763fe26d
We said no imaging study of peer comparison existed. Two had been published. We had not looked hard enough.
- Where
- Volume III and public evidence copy — the claim VI-E6 now guards against in lib/, app/, and components/.
- When
- 2026-08-13 (August 13, 2026)
- Why it was wrong
- Halpern 2021 (Journal of General Internal Medicine) and Clark-Randall 2021 (Journal of the American College of Radiology) had already reported observational pre-post imaging peer comparison in Duke Primary Care. One cohort, two denominators — volume and cost — not two independent findings. No concurrent control. The papers were in PubMed. We had not looked hard enough.
- What we say now
- Imaging-specific peer-comparison evidence exists and is uncontrolled. The open questions are how much of the published reduction was the intervention, whether imaging behaves like the antibiotic trials, and whether any effect persists — not whether a study had been published.
- Commit that fixed it
- e400d3ad504037d63ab4f0aa8b16dcbf7d20f949
We announced the ARKA-IP inpatient imaging benchmark as published on 18 August 2026 with a citation that pointed at a private repository and no DOI.
- Where
- docs/PUBLISHED_ASSETS.md ungated row for arka-ip-benchmark-v1.0; tools/arka-ip/benchmark/CITATION.cff identifiers (type: url into github.com/ArKa3003/arkahealth); the 18 August public announcement.
- When
- 2026-08-18 (August 18, 2026)
- Why it was wrong
- A dataset without a resolvable identifier is not published. A GitHub URL into a private monorepo cannot be cited, resolved, or retrieved. We marked the row published anyway.
- What we say now
- A dataset without a resolvable identifier is not published. The DOI was minted on 21 August 2026 (10.5281/zenodo.22051812). CHECK-BENCH-1 now fails any asset marked published without one.
- Commit that fixed it
- f809f97dd4531f61a86a8e3fe752b999b7f5b3b5
We treated evidence-quality abstention as a provenance filter: AIIE abstained only when GRADE certainty was very_low and the sole anchor was single-specialty AUC or payer-authored. Everything else, including strong recommendations on very-low certainty, could still emit a confident card.
- Where
- lib/evidence/grading.ts weakEvidenceMustAbstain; lib/aiie-v3/risk-control.ts evidenceQualityAbstain; docs/certainty-coverage.json abstainCount (was 0 against 1,349 very_low ratings).
- When
- 2026-08-22 (August 22, 2026)
- Why it was wrong
- Abstention is a function of certainty and consequence, not provenance alone. A strong recommendation on very-low certainty is the GRADE discordance case clinicians must not see as a confident card. The old rule caught almost nothing in the live matrix (single_specialty_auc and payer_utilization anchors are rare); the census stayed at abstainCount 0 while most very_low ratings were strong.
- What we say now
- evidenceQualityAbstain is the disjunction of weakEvidenceMustAbstain (unchanged) and strongRecOnVeryLowMustAbstain, which now fires only for a strong recommendation to image (rating ≥ 7) on very_low certainty. A strong recommendation not to image on very_low certainty still emits, and must carry formatClinicianCertaintyPhrase and STRONG_REC_LOW_CERTAINTY_CAUTION. Certainty values are not raised to shrink abstainCount — that is step 33.2 only.
- Commit that fixed it
- 7427e2fb88c8f4c3e5c17ee21b87e22c3679ba16
We claimed that nobody owns the per-clinician risk-adjusted ledger — that the measurement substrate for low-value imaging in risk-bearing primary care was unowned.
- Where
- Volume III prior positioning, pipeline qualification copy, and outreach materials — retired by the August 2026 four-layer prior-art scan (Prompt M4.4); public section on /evidence/nudges#prior-art-scan.
- When
- 2026-08-23 (August 23, 2026)
- Why it was wrong
- Milliman, Arcadia, Innovaccer, and Clarify each publish analytics that hold the D2K measurement substrate. The scan found the substrate is crowded; the unsettled layer is behavioural intervention above it, not ledger math.
- What we say now
- We no longer claim ledger exclusivity. We claim the intervention stack — locked pre-period, monthly peer packet outside the EHR, record-write verification — is what none of the named incumbents ship as a published product. The withdrawal is recorded beside the nudge write-up with sources that resolve.
- Commit that fixed it
- d3405a09a2e7543766df9dbfb4b473b2911293ad