AskAjay.ai
All tools

Instrument

Liability Ledger Audit

The Liability Ledger Audit scores an AI portfolio on 25 dimensions across five categories of ethical debt, from 25 (debt free) to 125 (critical). Each category compounds at its own rate every six months, and the audit puts the fastest-compounding debt at the top of the repair list.

Last verified

Scoring

The 25 dimensions, and what a score of 1, 3 and 5 looks like
Dimension1 Well-Managed3 Partial5 Critical
D1 Bias Debt: compounds 2.0x every six months
D1.1 Fairness testing coverageEvery user-facing model has had a fairness audit in the last six months, with automated tests in the deployment pipeline.30 to 60% of user-facing models audited, with no plan to cover the rest.No production model has ever been tested for fairness.
D1.2 Outcome monitoringAutomated monitoring of outcome gaps between groups on every high-risk model, with alerts and drift tracked since launch.One or two flagship models are monitored. The rest are not.Outcomes are not tracked by group. You could not produce disparate-impact data if a regulator asked.
D1.3 Proxy variable auditProxy risk analyzed for every high-risk model. Known proxies removed, or kept with a written reason and a control.Proxy analysis done only for the most visible model.Never done. "The model does not see race" is treated as proof.
D1.4 Remediation protocolA written bias playbook, tested in at least one real or simulated event.The data science team investigates ad hoc, with no roles or communication plan.No process and no precedent.
D1.5 Audit recencyA portfolio-wide audit within the past six months, with high-risk models audited more often.Last audit 12 to 18 months ago. The schedule slips.No fairness audit has ever been run.
D2 Transparency Debt: compounds 1.3x every six months
D2.1 Explainability coverageEvery customer-facing and high-risk model gives explanations suited to its audience, tested with users.Flagship or regulated models explain themselves. Most others are black boxes.No model can explain a decision.
D2.2 Model documentationA current model card for every production model, generated from the pipeline where possible.The central team documents its models. Business-unit and shadow AI go undocumented.No documentation. Losing one data scientist would leave models unmaintainable.
D2.3 Stakeholder communicationPeople are told when AI shapes a decision about them, in wording tested for comprehension.AI use appears in the privacy policy but not at the point of decision.AI use is hidden or never mentioned.
D2.4 Audit trailTamper-evident logs of input, model version and output for every high-risk decision, retrievable within 48 hours.Logs exist for some systems but miss inputs or versions. Retrieval is slow.Decisions are not logged.
D2.5 Regulatory readinessMeets current transparency rules in every jurisdiction, with work under way on the rules still to come.Meets the obvious requirements, with gaps elsewhere. Tracking is reactive.No one has assessed which transparency rules apply.
D3 Governance Debt: compounds 1.5x every six months
D3.1 Governance structureA cross-functional body with a charter, authority and a record of decisions, including stopping or retiring systems.A part-time or advisory function that cannot block a deployment.No structure. Whoever can deploy AI does.
D3.2 Policy coveragePolicies cover development, deployment, procurement and use. Technical controls enforce them, and they are reviewed every year.Partial, outdated or unenforced policies. Shadow AI is not covered.No AI policy of any kind.
D3.3 Review cadenceEvery production system reviewed at least quarterly, high-risk systems monthly or continuously.Annual reviews for some systems, none for others.Systems run indefinitely without review.
D3.4 Incident responseAn AI incident playbook with severity levels, roles and message templates, rehearsed.The general IT incident process, with nothing specific to AI.AI failures are not recognized as incidents.
D3.5 Accountability assignmentEvery production system has a named owner with authority and resources, kept through handovers.Informal owners, usually whoever built the system. No registry.Nobody owns any system’s outcomes.
D4 Privacy Debt: compounds 1.8x every six months
D4.1 Consent managementConsent covers AI use by purpose, and a withdrawal reaches the training pipeline.General consent that does not name AI use. A withdrawal stops at the source data.Data is used because it is available. No consent covers AI use.
D4.2 Data provenanceEvery training record traces to its source, consent status and permitted use.Sources are known for the main model. Consent status is mostly unknown.Training data origins are unknown.
D4.3 Cross-border complianceAI data flows are mapped and covered by transfer mechanisms in every jurisdiction.Requirements are known, mapping is incomplete and compliance is assumed.Data pooled across jurisdictions with no transfer assessment.
D4.4 Biometric data handlingMeets every applicable biometric law with specific written consent, or the organization has confirmed it processes none.Biometric use is documented, but general consent is treated as enough.Biometric data is processed with no awareness of biometric law.
D4.5 Retention and minimizationEnforced retention periods, regular minimization, and deletion requests that reach training sets and trigger retraining where needed.Retention rules are not applied to training data. Deletions stop at the source.Training data is kept indefinitely, and nobody can say what is in it.
D5 Accountability Debt: compounds 1.5x every six months
D5.1 Human oversightRisk-calibrated human review of every high-risk decision, by reviewers with the authority, time and information to override.Review on headline systems only, mostly approving what the model recommends.Consequential decisions are made with no human review.
D5.2 Escalation pathwaysAnyone affected can challenge an AI decision through a documented, tested path with timelines and a named decision-maker.Informal paths that depend on knowing whom to call.AI decisions are final.
D5.3 Vendor liability allocationVendor contracts carry AI-specific indemnity, audit rights, data-use limits and incident notice.Closer review than for ordinary software, but generic indemnity.Standard liability caps with no AI terms.
D5.4 Decision attributionLogs show whether the AI, a person or both made each decision, producible within 48 hours.Some logs, but AI and human decisions are not told apart.No record separates AI decisions from human ones.
D5.5 Redress mechanismsA documented, tested process to put right harm done to a person, with timelines and defined remedies.Handled case by case, depending on who picks up the complaint.No internal route to a remedy.

Score every dimension from 1 to 5 on evidence you can show. Scores 2 (Monitored) and 4 (Minimal) sit between the columns; the worksheet describes all five levels. For each dimension, keep the worst score across all your AI systems. Each category totals 5 to 25, and the five together give a score from 25 to 125. Lower is better.

Score bands
TotalBandWhat it means
25‑40Debt FreeDebt is measured and controlled, and monitoring is running. Keep the quarterly review.
41‑65Manageable DebtKnown debts with a paydown plan under way. Check that no single category is close to its maximum.
66‑90Dangerous DebtDebt is compounding unseen across several categories, and a lawsuit, inquiry or model failure would expose it. Start the 90-day plan now.
91‑125Critical DebtOne incident from crisis. Treat it as a board-level risk, with an executive sponsor and funding this quarter.
Subscriber Resource

Download: Liability Ledger Worksheet

The printable worksheet: every dimension, the score bands and room for your team's evidence.

Enter your email to get instant access. You'll also receive the fortnightly newsletter.

Free. No spam. Unsubscribe anytime.

Why this instrument exists, and what the scores mean for a leadership team: The Liability Ledger: How AI Liability Compounds

How to run the audit

  1. List every AI system in production, including vendor AI, shadow AI and internal tools. If the list cannot be finished, that gap is itself Governance Debt: score D3 accordingly.
  2. Score all 25 dimensions for each system on evidence. A planned audit, a draft policy or a committee that never meets scores as if it did not exist.
  3. For each dimension, keep the worst score across the portfolio. One critical system sets the ledger.
  4. Add each category, then the five categories, and read the band.
  5. Rank every dimension scored 4 or 5 by its category rate: Bias first, then Privacy, then Governance and Accountability, then Transparency.

Read the categories before the total

A total can look comfortable while one category compounds at a high rate. An organization at 60 overall sits in Manageable Debt, yet with Privacy Debt at 20 of 25 it carries debt that grows 1.8 times every six months. Build the plan from the dimensions scored 4 or 5, and use the band to report progress.

The multipliers below are directional. They come from enforcement patterns, settlement trajectories and regulatory timelines, since no actuarial data on ethical debt exists yet. Your sector and jurisdictions may run faster or slower. The order holds: Bias Debt compounds fastest and Transparency Debt slowest.

What waiting costs: remediation cost multipliers by category
CategoryRate per six monthsAfter 6 monthsAfter 12 monthsAfter 18 months
D1 Bias Debt2.0x2.0x4.0x8.0x
D4 Privacy Debt1.8x1.8x3.24x5.83x
D3 Governance Debt1.5x1.5x2.25x3.38x
D5 Accountability Debt1.5x1.5x2.25x3.38x
D2 Transparency Debt1.3x1.3x1.69x2.20x

Multiply what a fix costs today by the figure for its category and delay. A bias fix that costs $100,000 now costs about $800,000 after 18 months. The multipliers apply to cost only; the score stays on the 25 to 125 scale.

The 90-day reduction plan, ordered by compound rate
WeeksPhaseWhat the team doesOutput
1-2InventoryComplete the AI inventory and score all 25 dimensions. Start no fixes until the whole ledger is visible, so a large hidden debt does not lose out to a small visible one.Baseline score, five category subtotals, and the cost of waiting on every 4 and 5
3-4PrioritizeRank every 4 and 5 by category rate. Estimate what each fix costs today and set it against the 6 and 12 month multipliers. Secure an executive sponsor and a budget.An eight-week remediation plan with a named owner and a date for each item
5-8RemediateWork down the list. Bias first: fairness audits on the three highest-risk models, then outcome monitoring. Privacy next: data classification, consent gaps, cross-border flows. Then governance: a governance body, policies, a review cadence and an incident plan.Evidence for each fix, ready to re-score
9-12VerifyRe-score all 25 dimensions on the new evidence. Report the change in score and the cost of waiting that was avoided. Set a review cadence for each category, quarterly at the least.A board report and the date of the next quarterly review

Sources

  1. Temporal quality degradation in AI models (Vela et al., Scientific Reports 12, 11654). Nature Portfolio, 2022-07-08
  2. iTutorGroup to Pay $365,000 to Settle EEOC Discriminatory Hiring Suit. U.S. Equal Employment Opportunity Commission, 2023-09-11
  3. Attorney General Ken Paxton Secures $1.4 Billion Settlement with Meta Over Its Unauthorized Capture of Personal Biometric Data. Office of the Texas Attorney General, 2024-07-30
  4. AG Campbell Announces $2.5 Million Settlement With Student Loan Lender For Unlawful Practices Through AI Use. Massachusetts Office of the Attorney General, 2025-07-10
  5. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). EUR-Lex, Official Journal of the European Union, 2024-06-13
  6. AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 2023-01-26

Ajay's views, from 15 years in the field. Not legal or compliance advice. See full disclaimers →
Published by AI Exponent LLC

Liability Ledger Audit: Score AI Ethical Debt on 25 Dimensions | AskAjay.ai