Instrument
Liability Ledger Audit
The Liability Ledger Audit scores an AI portfolio on 25 dimensions across five categories of ethical debt, from 25 (debt free) to 125 (critical). Each category compounds at its own rate every six months, and the audit puts the fastest-compounding debt at the top of the repair list.
Last verified
Scoring
| Dimension | 1 Well-Managed | 3 Partial | 5 Critical |
|---|---|---|---|
| D1 Bias Debt: compounds 2.0x every six months | |||
| D1.1 Fairness testing coverage | Every user-facing model has had a fairness audit in the last six months, with automated tests in the deployment pipeline. | 30 to 60% of user-facing models audited, with no plan to cover the rest. | No production model has ever been tested for fairness. |
| D1.2 Outcome monitoring | Automated monitoring of outcome gaps between groups on every high-risk model, with alerts and drift tracked since launch. | One or two flagship models are monitored. The rest are not. | Outcomes are not tracked by group. You could not produce disparate-impact data if a regulator asked. |
| D1.3 Proxy variable audit | Proxy risk analyzed for every high-risk model. Known proxies removed, or kept with a written reason and a control. | Proxy analysis done only for the most visible model. | Never done. "The model does not see race" is treated as proof. |
| D1.4 Remediation protocol | A written bias playbook, tested in at least one real or simulated event. | The data science team investigates ad hoc, with no roles or communication plan. | No process and no precedent. |
| D1.5 Audit recency | A portfolio-wide audit within the past six months, with high-risk models audited more often. | Last audit 12 to 18 months ago. The schedule slips. | No fairness audit has ever been run. |
| D2 Transparency Debt: compounds 1.3x every six months | |||
| D2.1 Explainability coverage | Every customer-facing and high-risk model gives explanations suited to its audience, tested with users. | Flagship or regulated models explain themselves. Most others are black boxes. | No model can explain a decision. |
| D2.2 Model documentation | A current model card for every production model, generated from the pipeline where possible. | The central team documents its models. Business-unit and shadow AI go undocumented. | No documentation. Losing one data scientist would leave models unmaintainable. |
| D2.3 Stakeholder communication | People are told when AI shapes a decision about them, in wording tested for comprehension. | AI use appears in the privacy policy but not at the point of decision. | AI use is hidden or never mentioned. |
| D2.4 Audit trail | Tamper-evident logs of input, model version and output for every high-risk decision, retrievable within 48 hours. | Logs exist for some systems but miss inputs or versions. Retrieval is slow. | Decisions are not logged. |
| D2.5 Regulatory readiness | Meets current transparency rules in every jurisdiction, with work under way on the rules still to come. | Meets the obvious requirements, with gaps elsewhere. Tracking is reactive. | No one has assessed which transparency rules apply. |
| D3 Governance Debt: compounds 1.5x every six months | |||
| D3.1 Governance structure | A cross-functional body with a charter, authority and a record of decisions, including stopping or retiring systems. | A part-time or advisory function that cannot block a deployment. | No structure. Whoever can deploy AI does. |
| D3.2 Policy coverage | Policies cover development, deployment, procurement and use. Technical controls enforce them, and they are reviewed every year. | Partial, outdated or unenforced policies. Shadow AI is not covered. | No AI policy of any kind. |
| D3.3 Review cadence | Every production system reviewed at least quarterly, high-risk systems monthly or continuously. | Annual reviews for some systems, none for others. | Systems run indefinitely without review. |
| D3.4 Incident response | An AI incident playbook with severity levels, roles and message templates, rehearsed. | The general IT incident process, with nothing specific to AI. | AI failures are not recognized as incidents. |
| D3.5 Accountability assignment | Every production system has a named owner with authority and resources, kept through handovers. | Informal owners, usually whoever built the system. No registry. | Nobody owns any system’s outcomes. |
| D4 Privacy Debt: compounds 1.8x every six months | |||
| D4.1 Consent management | Consent covers AI use by purpose, and a withdrawal reaches the training pipeline. | General consent that does not name AI use. A withdrawal stops at the source data. | Data is used because it is available. No consent covers AI use. |
| D4.2 Data provenance | Every training record traces to its source, consent status and permitted use. | Sources are known for the main model. Consent status is mostly unknown. | Training data origins are unknown. |
| D4.3 Cross-border compliance | AI data flows are mapped and covered by transfer mechanisms in every jurisdiction. | Requirements are known, mapping is incomplete and compliance is assumed. | Data pooled across jurisdictions with no transfer assessment. |
| D4.4 Biometric data handling | Meets every applicable biometric law with specific written consent, or the organization has confirmed it processes none. | Biometric use is documented, but general consent is treated as enough. | Biometric data is processed with no awareness of biometric law. |
| D4.5 Retention and minimization | Enforced retention periods, regular minimization, and deletion requests that reach training sets and trigger retraining where needed. | Retention rules are not applied to training data. Deletions stop at the source. | Training data is kept indefinitely, and nobody can say what is in it. |
| D5 Accountability Debt: compounds 1.5x every six months | |||
| D5.1 Human oversight | Risk-calibrated human review of every high-risk decision, by reviewers with the authority, time and information to override. | Review on headline systems only, mostly approving what the model recommends. | Consequential decisions are made with no human review. |
| D5.2 Escalation pathways | Anyone affected can challenge an AI decision through a documented, tested path with timelines and a named decision-maker. | Informal paths that depend on knowing whom to call. | AI decisions are final. |
| D5.3 Vendor liability allocation | Vendor contracts carry AI-specific indemnity, audit rights, data-use limits and incident notice. | Closer review than for ordinary software, but generic indemnity. | Standard liability caps with no AI terms. |
| D5.4 Decision attribution | Logs show whether the AI, a person or both made each decision, producible within 48 hours. | Some logs, but AI and human decisions are not told apart. | No record separates AI decisions from human ones. |
| D5.5 Redress mechanisms | A documented, tested process to put right harm done to a person, with timelines and defined remedies. | Handled case by case, depending on who picks up the complaint. | No internal route to a remedy. |
Score every dimension from 1 to 5 on evidence you can show. Scores 2 (Monitored) and 4 (Minimal) sit between the columns; the worksheet describes all five levels. For each dimension, keep the worst score across all your AI systems. Each category totals 5 to 25, and the five together give a score from 25 to 125. Lower is better.
| Total | Band | What it means |
|---|---|---|
| 25‑40 | Debt Free | Debt is measured and controlled, and monitoring is running. Keep the quarterly review. |
| 41‑65 | Manageable Debt | Known debts with a paydown plan under way. Check that no single category is close to its maximum. |
| 66‑90 | Dangerous Debt | Debt is compounding unseen across several categories, and a lawsuit, inquiry or model failure would expose it. Start the 90-day plan now. |
| 91‑125 | Critical Debt | One incident from crisis. Treat it as a board-level risk, with an executive sponsor and funding this quarter. |
Download: Liability Ledger Worksheet
The printable worksheet: every dimension, the score bands and room for your team's evidence.
Enter your email to get instant access. You'll also receive the fortnightly newsletter.
Free. No spam. Unsubscribe anytime.
Why this instrument exists, and what the scores mean for a leadership team: The Liability Ledger: How AI Liability Compounds
How to run the audit
- List every AI system in production, including vendor AI, shadow AI and internal tools. If the list cannot be finished, that gap is itself Governance Debt: score D3 accordingly.
- Score all 25 dimensions for each system on evidence. A planned audit, a draft policy or a committee that never meets scores as if it did not exist.
- For each dimension, keep the worst score across the portfolio. One critical system sets the ledger.
- Add each category, then the five categories, and read the band.
- Rank every dimension scored 4 or 5 by its category rate: Bias first, then Privacy, then Governance and Accountability, then Transparency.
Read the categories before the total
A total can look comfortable while one category compounds at a high rate. An organization at 60 overall sits in Manageable Debt, yet with Privacy Debt at 20 of 25 it carries debt that grows 1.8 times every six months. Build the plan from the dimensions scored 4 or 5, and use the band to report progress.
The multipliers below are directional. They come from enforcement patterns, settlement trajectories and regulatory timelines, since no actuarial data on ethical debt exists yet. Your sector and jurisdictions may run faster or slower. The order holds: Bias Debt compounds fastest and Transparency Debt slowest.
| Category | Rate per six months | After 6 months | After 12 months | After 18 months |
|---|---|---|---|---|
| D1 Bias Debt | 2.0x | 2.0x | 4.0x | 8.0x |
| D4 Privacy Debt | 1.8x | 1.8x | 3.24x | 5.83x |
| D3 Governance Debt | 1.5x | 1.5x | 2.25x | 3.38x |
| D5 Accountability Debt | 1.5x | 1.5x | 2.25x | 3.38x |
| D2 Transparency Debt | 1.3x | 1.3x | 1.69x | 2.20x |
Multiply what a fix costs today by the figure for its category and delay. A bias fix that costs $100,000 now costs about $800,000 after 18 months. The multipliers apply to cost only; the score stays on the 25 to 125 scale.
| Weeks | Phase | What the team does | Output |
|---|---|---|---|
| 1-2 | Inventory | Complete the AI inventory and score all 25 dimensions. Start no fixes until the whole ledger is visible, so a large hidden debt does not lose out to a small visible one. | Baseline score, five category subtotals, and the cost of waiting on every 4 and 5 |
| 3-4 | Prioritize | Rank every 4 and 5 by category rate. Estimate what each fix costs today and set it against the 6 and 12 month multipliers. Secure an executive sponsor and a budget. | An eight-week remediation plan with a named owner and a date for each item |
| 5-8 | Remediate | Work down the list. Bias first: fairness audits on the three highest-risk models, then outcome monitoring. Privacy next: data classification, consent gaps, cross-border flows. Then governance: a governance body, policies, a review cadence and an incident plan. | Evidence for each fix, ready to re-score |
| 9-12 | Verify | Re-score all 25 dimensions on the new evidence. Report the change in score and the cost of waiting that was avoided. Set a review cadence for each category, quarterly at the least. | A board report and the date of the next quarterly review |
Sources
- Temporal quality degradation in AI models (Vela et al., Scientific Reports 12, 11654). Nature Portfolio, 2022-07-08
- iTutorGroup to Pay $365,000 to Settle EEOC Discriminatory Hiring Suit. U.S. Equal Employment Opportunity Commission, 2023-09-11
- Attorney General Ken Paxton Secures $1.4 Billion Settlement with Meta Over Its Unauthorized Capture of Personal Biometric Data. Office of the Texas Attorney General, 2024-07-30
- AG Campbell Announces $2.5 Million Settlement With Student Loan Lender For Unlawful Practices Through AI Use. Massachusetts Office of the Attorney General, 2025-07-10
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). EUR-Lex, Official Journal of the European Union, 2024-06-13
- AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 2023-01-26
Ajay's views, from 15 years in the field. Not legal or compliance advice. See full disclaimers →
Published by AI Exponent LLC