Inventorying, scoring, and reducing the hidden liability accumulating in your AI portfolio. Like a financial audit surfaces hidden debt, the Liability Ledger surfaces ethical debt — and prices the cost of delay.
Canonical references: Article 1 — How AI Liability Compounds · Article 2 — The Measurement & Reduction Framework
Liability compounds. Governance pays it down. The five debt categories accrue interest at different rates — Bias at 2.0x per six months, Privacy at 1.8x, Governance and Accountability at 1.5x, Transparency at 1.3x. Compounding is not a metaphor: every quarter you defer remediation, the cost of remediation rises measurably.
Directional, not actuarial. The Liability Ledger is a diagnostic tool, not an insurance underwriting model. It tells you where to look, how to prioritize, and what the cost of delay is likely to be. It does not tell you the cost to the penny. The full Limitations and Counter-Evidence sections appear at the end of this worksheet — read them before presenting scores externally.
A framework that hides its limitations is a framework you should not trust. Three caveats apply to every score this worksheet produces:
List all AI systems currently in production, including shadow AI and third-party AI tools. Inventory first, prioritize second — do not begin scoring until the full picture is visible.
| System Name | Owner | Production Since | Users Affected | Risk Tier | Last Audit |
|---|---|---|---|---|---|
| | |||||
| | |||||
| | |||||
| | |||||
| | |||||
| | |||||
| | |||||
| |
Untested models making decisions about people without fairness audits. Every quarter without a fairness audit is a quarter where the legal definition of "reasonable care" is tightening while your exposure remains unaddressed. Bias Debt held for 18 months costs approximately 8x what it would have cost to address at deployment.
| 1 | Well-Managed | All user-facing models have completed fairness audits within the past 6 months. Testing covers all legally protected characteristics relevant to the use case. Automated fairness testing is integrated into the deployment pipeline. |
| 2 | Monitored | 60-90% of user-facing models audited within the past 12 months. Known gaps are documented with scheduled remediation dates. Fairness testing exists but is not yet automated. |
| 3 | Partial | 30-60% of user-facing models audited. Testing is inconsistent — some models audited rigorously, others not at all. No systematic coverage plan. |
| 4 | Minimal | Less than 30% of user-facing models audited. Fairness testing is ad hoc, triggered by incidents or media attention rather than policy. Most models in production have never been tested. |
| 5 | Critical | No fairness testing conducted on any production model. No awareness of which models make decisions about protected characteristics. No testing methodology or tooling in place. |
| 1 | Well-Managed | Continuous automated monitoring for disparate impact across all high-risk models in production. Alerts trigger review when differential outcomes cross defined thresholds. Dashboard accessible to governance team. Drift since deployment is tracked, not just point-in-time fairness at launch. |
| 2 | Monitored | Regular monitoring (monthly/quarterly) for key models. Thresholds defined for most protected characteristics. Manual review process in place. Some gaps in coverage; drift detection exists but is not real-time. |
| 3 | Partial | Monitoring exists for one or two flagship models, typically those with the highest external visibility. Most models operate without disparate-impact tracking. Drift between deployment-time fairness and current production outcomes is not measured. |
| 4 | Minimal | No systematic monitoring. Disparate impact is detected only through customer complaints, lawsuits, or media attention. Reactive, not proactive. The organization treats deployment-time fairness testing as sufficient. |
| 5 | Critical | No concept of outcome monitoring. Outcomes are not tracked by population segment. The organization could not produce disparate-impact data even if a regulator required it. A model that passed fairness testing at deployment could be producing discriminatory outcomes today and no one would know. |
| 1 | Well-Managed | Documented analysis of proxy discrimination risk for all high-risk models. Known proxies (zip code, name, school, employer history, device type) either removed or retained only with explicit, documented justification and compensating controls. Proxy audit is a deployment-gate requirement. |
| 2 | Monitored | Proxy variable analysis conducted for new high-risk models and being retro-applied to existing ones. Common proxies catalogued. Some gaps in feature engineering review for older systems. Justification for retained proxies is being formalized. |
| 3 | Partial | Proxy risk understood as a concept; analysis performed only for the highest-profile model. Most teams treat "we removed protected attributes" as sufficient. The Earnest / HBCU pattern (race-neutral inputs producing race-correlated outputs) has not been systematically tested for. |
| 4 | Minimal | Vague awareness that proxy variables exist. No analysis performed. Feature selection driven by predictive accuracy alone. The team would not be able to name the top three proxy risks in their flagship model. |
| 5 | Critical | No proxy variable analysis has ever been conducted. The concept is not part of the model development process. Leadership treats "the model does not see race" as proof the model cannot discriminate. This is the most common bias-debt failure mode in production AI today. |
| 1 | Well-Managed | Documented remediation playbook covering identification, triage, root cause analysis, remediation options, stakeholder notification, implementation, and verification. Playbook has been tested through at least one real or simulated event. |
| 2 | Monitored | Remediation process documented but not yet tested. Key steps defined: who to notify, how to investigate, when to pause a model. Process gaps acknowledged with improvement plans. |
| 3 | Partial | Informal remediation approach. When bias is found, the data science team investigates ad hoc. No documented process, no defined roles, no communication plan. |
| 4 | Minimal | No remediation process. Bias findings have been acknowledged but not acted upon. "We will fix it in the next model version" is the default response. |
| 5 | Critical | No process and no precedent. Bias has never been investigated because it has never been tested for. If bias were discovered tomorrow, the organization would not know where to begin. |
| 1 | Well-Managed | Most recent portfolio-wide fairness audit completed within the past 6 months. Individual high-risk models audited more frequently (quarterly or continuously). Audit cadence aligned with model drift rates. |
| 2 | Monitored | Most recent portfolio-wide audit within the past 12 months. High-risk models audited within 6 months. Audit schedule documented and followed. |
| 3 | Partial | Most recent audit 12-18 months ago. Some models audited more recently due to updates or incidents, but no systematic cadence. Audit schedule exists on paper but slips regularly. |
| 4 | Minimal | Most recent audit over 18 months ago, or only a subset of models have ever been audited. No recurring audit cadence. |
| 5 | Critical | No fairness audit has ever been conducted. The concept of an AI fairness audit is unfamiliar to the organization's leadership. |
Black-box systems with no explainability for stakeholders. Models making consequential decisions that no one can explain or interrogate. The EU AI Act's high-risk-system transparency obligations carry penalties of up to EUR 15 million or 3% of global annual turnover under Article 99(4). (The 7% / EUR 35 million tier under Article 99(3) applies only to prohibited practices — not to transparency violations.)
| 1 | Well-Managed | All customer-facing and high-risk models produce audience-appropriate explanations. Explainability tested with actual users. Multiple explanation formats available (technical, business, plain language). |
| 2 | Monitored | 60-90% of high-risk models have explainability capability. Explanations are available but may not be tailored to audience. Coverage plan exists for remaining systems. |
| 3 | Partial | Explainability exists for flagship models or those with regulatory requirements, but most models operate as black boxes. No systematic explainability standard. |
| 4 | Minimal | Explainability exists for one or two models, typically developed in response to a regulatory inquiry or customer complaint. Most systems have no explanation capability. |
| 5 | Critical | No model in the portfolio can produce a human-readable explanation of its decisions. "The model said so" is the only available rationale. |
| 1 | Well-Managed | All production models have comprehensive, current documentation (model cards or equivalent). Documentation is auto-generated from the ML pipeline where possible. Includes purpose, training data, performance metrics, limitations, ownership, and update history. |
| 2 | Monitored | 60-90% of production models documented. Documentation follows a standard template. Review cadence exists, though some documentation may be slightly outdated. |
| 3 | Partial | Some models documented, typically those built by the central data science team. Business-unit-deployed models and shadow AI lack documentation. No standard template enforced. |
| 4 | Minimal | Documentation exists for one or two models, created for a specific compliance requirement. Most models have no documentation. The organization does not know how many AI models are in production. |
| 5 | Critical | No model documentation exists. Training data provenance unknown. Decision logic undocumented. The departure of a key data scientist would render models unmaintainable. |
| 1 | Well-Managed | Proactive, clear communication about AI use in all relevant contexts. Customers can access explanations of AI-influenced decisions. Communication tested with actual users for comprehension. |
| 2 | Monitored | AI use disclosed in privacy policies and relevant product documentation. Customers informed at key decision points. Some explanation capability available on request. |
| 3 | Partial | AI use mentioned in privacy policy but not at the point of decision. Most customers unaware that AI influences their experience. Explanations available only if specifically demanded. |
| 4 | Minimal | AI use not disclosed to customers. Internal stakeholders have limited awareness of which processes use AI. Transparency is treated as a liability rather than a practice. |
| 5 | Critical | Active concealment of AI use, or complete absence of any communication. Customers would be surprised to learn AI is involved in decisions affecting them. |
| 1 | Well-Managed | Complete, tamper-evident audit trail for all high-risk AI decisions. Decision inputs, model version, confidence scores, and outputs logged and retained. Trails queryable for regulatory response within 48 hours. |
| 2 | Monitored | Audit trails exist for most high-risk systems. Logging covers key decision attributes. Some gaps in older systems or edge cases. Retention policies defined. |
| 3 | Partial | Audit trails exist for some systems but are incomplete or inconsistent. Logging may capture outputs but not inputs or model versions. Retrieval is possible but slow. |
| 4 | Minimal | Limited logging, primarily for debugging rather than governance. Logs are not designed for regulatory or legal purposes. Retention is ad hoc. Retrieving a specific decision history is difficult or impossible. |
| 5 | Critical | No audit trail. AI decisions are not logged. If a regulator or court requested the basis for a specific AI decision, the organization could not produce it. |
| 1 | Well-Managed | Full compliance with current transparency regulations across all applicable jurisdictions. Proactive preparation for forthcoming requirements. Explainability infrastructure designed to scale with regulatory evolution. |
| 2 | Monitored | Compliance with most current requirements. Gaps identified and scheduled for remediation. Regulatory tracking process in place. Some preparation for forthcoming requirements. |
| 3 | Partial | Compliance with the most immediate or obvious regulatory requirements, but gaps in less visible mandates. Limited awareness of forthcoming requirements. Reactive regulatory tracking. |
| 4 | Minimal | Non-compliance with one or more applicable transparency regulations. Awareness of regulatory requirements exists but no compliance plan. Transparency is treated as a future problem. |
| 5 | Critical | No awareness of applicable transparency regulation. No compliance assessment conducted. The organization would fail a regulatory audit on transparency grounds immediately. |
No oversight structure, no accountability, no review cadence. Shadow AI proliferating faster than sanctioned AI. IBM's 2025 Cost of a Data Breach Report found shadow AI breaches carry a $670K cost premium. Governance debt compounds because shadow AI proliferates exponentially in the absence of structure.
| 1 | Well-Managed | Cross-functional AI governance body with clear charter, defined authority, regular cadence, executive sponsorship, and demonstrated track record of decisions (including decisions to not deploy or to retire systems). |
| 2 | Monitored | Governance body exists with defined authority but is still maturing. Meets regularly. Has made governance decisions but may not yet cover all AI systems. Some gaps in scope or authority. |
| 3 | Partial | Governance function exists but lacks authority or consistency. May be one person's part-time role. Meets irregularly. Advisory capacity only — cannot block deployments. |
| 4 | Minimal | Governance exists on paper. A policy was written, a committee was named, but neither functions. No meetings held in the past 6 months. No decisions made. |
| 5 | Critical | No governance structure exists. No policy, no committee, no designated ownership. AI is deployed by whoever has the technical capability to do so. |
| 1 | Well-Managed | Comprehensive AI policy suite covering development, deployment, procurement, acceptable use, data handling, and third-party AI. Policies enforced through technical controls and governance processes. Reviewed and updated at least annually. |
| 2 | Monitored | Policies exist and cover major areas. Enforcement is active but relies partly on manual processes. Some policy gaps in emerging areas (GenAI acceptable use, agentic AI, third-party model procurement). |
| 3 | Partial | Some policies exist but coverage is incomplete. Enforcement is inconsistent. Policies may be outdated or not reflect current AI use patterns. Shadow AI is not addressed. |
| 4 | Minimal | A single AI ethics policy exists, drafted for a compliance requirement, and has not been updated or enforced since creation. |
| 5 | Critical | No AI policies of any kind. No acceptable use guidance. No development standards. No procurement criteria. Employees using AI make their own judgments about appropriate use. |
| 1 | Well-Managed | All production AI systems reviewed at least quarterly. High-risk systems reviewed monthly or continuously. Review covers performance, fairness, drift, compliance, and business alignment. |
| 2 | Monitored | Quarterly reviews for high-risk systems. Annual reviews for lower-risk systems. Review process standardized with checklists. Some reviews produce action items that are tracked. |
| 3 | Partial | Annual reviews for some systems. No regular cadence for others. Review depth varies. Some reviews are pro forma — a check-the-box exercise rather than a substantive evaluation. |
| 4 | Minimal | Reviews happen only when triggered by incidents, complaints, or regulatory inquiries. No proactive review cadence. Most systems have not been reviewed since initial deployment. |
| 5 | Critical | No review process exists. AI systems deployed to production operate indefinitely without evaluation. No mechanism to identify systems that have drifted, degraded, or become non-compliant. |
| 1 | Well-Managed | Documented AI incident response playbook with severity classification, assigned roles, escalation paths, communication templates, and remediation procedures. Tested through tabletop exercises or real incidents. |
| 2 | Monitored | AI incident response plan exists and has been communicated. Roles assigned. Severity levels defined. Not yet tested through a full simulation. Some process gaps acknowledged. |
| 3 | Partial | General IT incident response process exists but is not AI-specific. AI incidents are handled through the general process without AI-tailored procedures. |
| 4 | Minimal | No documented incident response for AI. Incidents handled ad hoc by whoever discovers them. No severity classification, no escalation path, no communication protocol. |
| 5 | Critical | No incident response capability and no awareness that AI-specific incidents require distinct handling. AI failures are not recognized as a category of incident. |
| 1 | Well-Managed | Every production AI system has a named human owner with documented accountability for outcomes, performance, and compliance. Ownership is maintained through transitions. Owners have authority and resources to fulfill their accountability. |
| 2 | Monitored | Most AI systems have named owners. Ownership documented in a central registry. Some gaps in legacy or recently deployed systems. Ownership accountability defined but not yet fully operational. |
| 3 | Partial | Some systems have named owners, typically the person who built them. Ownership is informal — based on who built it, not on a formal assignment. No central registry. |
| 4 | Minimal | Ownership is attributed to teams or departments rather than individuals. "The data science team owns it" is the standard answer. No individual feels personally accountable. |
| 5 | Critical | No ownership concept for AI systems. No one is responsible for any AI system's outcomes. Models run in production with no identifiable steward. |
Data used in AI systems without proper consent, classification, controls, or retention management. Meta's $1.4 billion Texas settlement and Clearview AI's $50 million fine demonstrate that privacy penalties are escalating rapidly. Privacy violations tend to be systemic — if one practice is non-compliant, similar practices across the organization likely are too.
| 1 | Well-Managed | Consent specifically addresses AI use. Granular consent management system tracks which data can be used for which AI purposes. Consent withdrawal triggers are connected to the ML pipeline (model retraining, data deletion, or both). Legacy consent has been re-papered for AI use. |
| 2 | Monitored | AI-specific consent is obtained for new data collection. Legacy data consent is being reviewed and remediated. Consent withdrawal process exists but may involve manual steps. Most data flows have explicit AI provisions. |
| 3 | Partial | General data collection consent exists but does not specifically address AI use. Consent may not cover the specific purposes for which data is being used in AI. No process to propagate consent withdrawal into AI systems. |
| 4 | Minimal | Consent management is rudimentary. Data is used in AI based on Terms of Service provisions rather than explicit, informed AI-specific consent. Significant uncertainty about whether existing consent covers current AI use. |
| 5 | Critical | No consent mechanism addresses AI use. Data is used because it is available, not because consent covers the use. Customer data has been uploaded to third-party AI tools without consent review. Broad, ambiguous terms are treated as permission to train. |
| 1 | Well-Managed | Complete provenance chains for all AI training data. Every record traceable to source, consent status, permissible uses, and date of acquisition. The organization could answer a regulator's question about training-data sourcing within hours, not weeks. |
| 2 | Monitored | Provenance documented for most training datasets. Source, consent, and permissible-use metadata captured for new acquisitions. Legacy datasets being back-filled. Some gaps in third-party datasets and synthetic data lineage. |
| 3 | Partial | Provenance is documented for the highest-profile model only. Most training data has known sources but unknown consent status. Permissible-use restrictions are not tracked at the record level. A regulator's data-sourcing question would require investigation. |
| 4 | Minimal | Provenance tracked at dataset level but not record level. Many datasets have unknown or partially-known origins. Web-scraped, third-party-licensed, and internally-merged data are commingled without lineage. The organization could not separate compliant from non-compliant data on demand. |
| 5 | Critical | Training data origins are unknown. The organization could not answer a regulator's question about data sourcing. Datasets used to train production models include data of unknown provenance. A consent-revocation request cannot be honored because the affected records cannot be located. |
| 1 | Well-Managed | Cross-border data flows for AI mapped and compliant across all jurisdictions where the organization operates. Transfer mechanisms (SCCs, adequacy decisions, BCRs) in place and current. Data residency requirements met for AI processing. Cloud AI services are gated by residency policy. |
| 2 | Monitored | Cross-border data flows for AI mostly mapped and compliant. Transfer mechanisms in place for major jurisdictions. Some gaps in newer data flows or smaller jurisdictions. Remediation underway. Cloud AI service residency reviewed before contracting. |
| 3 | Partial | Awareness of cross-border requirements exists, but mapping is incomplete. Compliance is assumed rather than verified. Transfer mechanisms may cover general data but not specifically address AI processing flows. |
| 4 | Minimal | Limited awareness of cross-border data implications for AI. Data is processed where the AI infrastructure is located without consideration of data origin. Cloud AI services used without data residency review. |
| 5 | Critical | No awareness of cross-border data compliance for AI. Data from multiple jurisdictions pooled without any cross-border assessment. AI training data likely includes data from jurisdictions with strict sovereignty requirements (EU, China, India, UAE) with no compliance measures in place. |
| 1 | Well-Managed | Full compliance with all applicable biometric laws (Illinois BIPA, Texas CUBI, Washington's biometric identifier law, EU AI Act biometric categorization rules). Explicit, written, biometric-specific consent obtained before processing. Biometric data segregated and access-controlled. Or: biometric data is not processed at all and the organization has confirmed the negative. |
| 2 | Monitored | Biometric processing inventoried and major jurisdictions covered. Biometric-specific consent obtained for new collection. Legacy biometric data being remediated. Some emerging-jurisdiction gaps acknowledged with closure plans. |
| 3 | Partial | Biometric processing exists and is documented at the system level, but biometric-specific compliance posture is unclear. General privacy consent is treated as covering biometric use. The Meta $1.4B Texas settlement is known but not used to drive a compliance review. |
| 4 | Minimal | Biometric data is processed (face matching, voice prints, gait, behavioral biometrics) but the organization has not reviewed biometric-specific obligations. No biometric-specific consent. Biometric data handled under generic data-protection policies. |
| 5 | Critical | Biometric data is processed without awareness of biometric-specific regulation. No inventory of which systems use biometric features. Per-violation statutory damages (BIPA: $1,000–$5,000 per violation) would multiply across the user base into a Meta-scale exposure. |
| 1 | Well-Managed | Documented AI-data retention policies with automated enforcement. Retention periods aligned with purpose and legal requirements. Regular minimization reviews remove unnecessary fields from active training sets. Deletion requests trigger verifiable removal from training sets and model retraining where necessary. |
| 2 | Monitored | Retention and minimization policies exist and cover AI data. Periodic reviews flag over-collection. Deletion process defined but may involve manual steps. Most data subject to defined retention periods. Some legacy data outside the framework. |
| 3 | Partial | General retention policies exist but are not consistently applied to AI training data. Minimization is understood but inconsistently practiced — "more data is better" still operates as a default. Deletion requests honored for source data but not propagated to AI training sets. |
| 4 | Minimal | Retention policies exist on paper but are not enforced for AI data. "We need the data for model training" overrides retention limits. No minimization discipline — models trained on every available field. Deletion requests are technically infeasible to propagate to AI systems. |
| 5 | Critical | No retention policies for AI data. Training data is accumulated indefinitely without review. Personal data treated as a raw material to be maximized. The organization cannot determine what data is in its AI training sets, let alone honor a deletion request. |
No one owns AI outcomes. No audit trail links a decision to the person who approved the system. The 2025 Workday ruling established that AI vendors can be directly liable for discriminatory outcomes. The legal definition of "accountable party" is expanding, and organizations without clear accountability chains are accumulating exposure.
| 1 | Well-Managed | Risk-calibrated human oversight for all high-risk AI decisions. Oversight is meaningful — the human reviewer has authority to override, with the time, training, and information needed to do so. Oversight design documented per system. EU AI Act Article 14 requirements demonstrably met for high-risk systems. |
| 2 | Monitored | Human oversight present for most high-risk AI decisions. Reviewers have override authority. Some systems use rubber-stamp review (high approval rates with low review time) and are being remediated. Oversight design improving but not yet uniform. |
| 3 | Partial | Human-in-the-loop exists for headline systems. Many production AI systems operate with minimal human review — the human approves what the model recommends. Reviewer training is informal; override rates are low and not tracked. |
| 4 | Minimal | Human review is nominal — humans see AI outputs but lack the authority, time, or information to meaningfully intervene. "Automation bias" is the operating reality. No documented oversight design per system. |
| 5 | Critical | AI makes consequential decisions autonomously with no human review at any stage. No defined override mechanism. The organization could not produce evidence of human oversight for a regulator or court. |
| 1 | Well-Managed | Documented escalation paths for any party (employee, customer, applicant, regulator) to challenge an AI decision. Defined timelines and named decision authority. Paths tested through real or simulated cases. Outcome data from challenged decisions feeds back into model improvement. |
| 2 | Monitored | Escalation paths defined for high-stakes AI decisions (employment, lending, benefits). Lower-stakes escalation relies on general support channels. Paths have been used at least once. Some gaps in coverage for non-customer-facing AI systems. |
| 3 | Partial | Escalation paths exist informally. People know who to contact if something goes wrong, but it depends on relationships rather than process. No published mechanism for affected parties to challenge an AI decision. No guaranteed timeframes. |
| 4 | Minimal | No AI-specific escalation. Concerns raised through general channels (helpdesk, customer service) with no guarantee of reaching someone who understands AI-specific implications or has the authority to overturn an AI decision. |
| 5 | Critical | No escalation mechanism. AI decisions are treated as final. Affected parties have no internal channel to challenge an outcome. Regulators expect a contestability surface; the organization has none. |
| 1 | Well-Managed | AI vendor contracts include explicit AI-specific liability terms: indemnification for algorithmic harms, audit rights, data-use restrictions, model-update notice requirements, and incident-disclosure obligations. Contracts have been re-papered post-Workday to anticipate third-party liability exposure. |
| 2 | Monitored | AI-specific liability terms in new contracts. Existing contracts being renegotiated as they renew. Audit rights and indemnification negotiated for high-risk vendors. Some legacy contracts still rely on standard limitation-of-liability caps. |
| 3 | Partial | Contract review for AI vendors is more rigorous than for general SaaS, but AI-specific provisions are inconsistent. Indemnification language is generic. Audit rights exist for some vendors but have never been exercised. |
| 4 | Minimal | AI vendor contracts use standard SaaS templates. Liability caps at contract value. No AI-specific indemnification. Vendors' model behavior, training data, and update cadence are not contractually visible. |
| 5 | Critical | Standard limitation-of-liability caps with no AI-specific provisions. The organization is exposed to algorithmic-harm claims with no contractual recourse to the vendor. The Workday precedent suggests this exposure now extends to third parties (job applicants, etc.) who never signed the contract at all. |
| 1 | Well-Managed | Clear attribution for every AI-influenced decision. Logs distinguish AI-driven, human-driven, and AI-recommended-then-human-confirmed outcomes. Model version, input features, confidence score, and human reviewer (if any) all captured. Attribution is producible for any decision within 48 hours. |
| 2 | Monitored | Attribution captured for high-risk systems. Most AI-influenced decisions distinguishable from purely human ones. Some gaps in legacy systems where attribution metadata was not built into the decision record. Producibility tested but may take longer than 48 hours. |
| 3 | Partial | Some logging exists but does not consistently distinguish AI-driven from human-driven outcomes. Model versions are not tied to specific decisions. After-the-fact reconstruction of who or what made a given decision is possible but slow. |
| 4 | Minimal | Limited attribution. Decisions are recorded as outcomes without capturing whether AI, human, or interaction produced them. "The system decided" is the operating phrase — it obscures rather than reveals. |
| 5 | Critical | No distinction between AI-driven and human-driven decisions in any record. Accountability is impossible to assign. A regulator, court, or affected individual asking "who made this decision?" cannot be given a defensible answer. |
| 1 | Well-Managed | Documented redress process for individuals harmed by AI decisions. Defined timelines, named decision authority, remedies appropriate to harm category (decision reversal, monetary compensation, model retraining, public correction). Process tested with real cases. Redress data feeds back into governance. |
| 2 | Monitored | Redress process exists for high-stakes AI decisions (employment, lending, benefits). Timelines defined; remedies catalogued. Process has been used at least once. Some gaps in coverage for less visible AI systems. |
| 3 | Partial | Redress is handled ad hoc, case by case. No documented process. Remedies depend on who handles the complaint and how visible the harm becomes. Affected individuals do not know what to expect. |
| 4 | Minimal | No formal redress mechanism. Complaints route through general customer service or legal. Remedy depends on litigation risk, not on harm severity. Most harmed individuals never reach a meaningful remedy. |
| 5 | Critical | No redress mechanism at all. Individuals harmed by AI have no internal avenue for remedy. Their only path is regulatory complaint, litigation, or public exposure — each of which converts the original harm into a much larger liability event for the organization. |
Category imbalance reveals the highest-interest debt. Prioritize the highest-interest debt, not the highest absolute score. An organization scoring 8 on D1 (Bias) and 22 on D2 (Transparency) should not average the scores — the D2 score is compounding at 1.3x while D1 is well-managed.
| Priority | Category | Compound Rate | Your Score | Severity |
|---|---|---|---|---|
| 1 | D1: Bias Debt | 2.0x / 6 months | ||
| 2 | D4: Privacy Debt | 1.8x / 6 months | ||
| 3 | D3: Governance Debt | 1.5x / 6 months | ||
| 4 | D5: Accountability Debt | 1.5x / 6 months | ||
| 5 | D2: Transparency Debt | 1.3x / 6 months |
For every 6 months that debt is held without action, compound interest multiplies the cost of remediation. Bias Debt held for 18 months costs approximately 8x what it would have cost to address at deployment.
| Category | Current Score | 6-Mo Rate | Projected 6 Mo | Projected 12 Mo | Projected 18 Mo |
|---|---|---|---|---|---|
| D1: Bias Debt | 2.0x | ||||
| D4: Privacy Debt | 1.8x | ||||
| D3: Governance Debt | 1.5x | ||||
| D5: Accountability Debt | 1.5x | ||||
| D2: Transparency Debt | 1.3x |
"Pay Down Highest Interest First." Borrowed from personal finance: when carrying multiple debts, pay down the highest interest rate first. This means Bias Debt first (2.0x), then Privacy Debt (1.8x), then Governance and Accountability Debt (1.5x each), then Transparency Debt (1.3x).
| Priority | Category | Current Score | Target Score | Actions Required | Owner | Deadline |
|---|---|---|---|---|---|---|
| 1 | | |||||
| 2 | | |||||
| 3 | | |||||
| 4 | | |||||
| 5 | |
Liability exposure varies by industry. A score of 60 may represent manageable debt in consumer technology but dangerous exposure in healthcare. Use these benchmarks to contextualize your score.
Liability sensitivity: Very High | Regulatory pressure: Very High | Litigation exposure: Extreme
| Category | Debt Free (<) | Manageable | Dangerous | Critical (>) |
|---|---|---|---|---|
| D1: Bias | <6 | 6–12 | 13–18 | >18 |
| D2: Transparency | <7 | 7–13 | 14–19 | >19 |
| D3: Governance | <6 | 6–11 | 12–17 | >17 |
| D4: Privacy | <6 | 6–12 | 13–18 | >18 |
| D5: Accountability | <6 | 6–11 | 12–17 | >17 |
| Total | <31 | 31–59 | 60–89 | >89 |
Liability sensitivity: Critical (life-safety) | Regulatory pressure: Very High | Litigation exposure: Very High
| Category | Debt Free (<) | Manageable | Dangerous | Critical (>) |
|---|---|---|---|---|
| D1: Bias | <6 | 6–11 | 12–17 | >17 |
| D2: Transparency | <6 | 6–12 | 13–18 | >18 |
| D3: Governance | <6 | 6–11 | 12–17 | >17 |
| D4: Privacy | <5 | 5–10 | 11–16 | >16 |
| D5: Accountability | <6 | 6–11 | 12–17 | >17 |
| Total | <29 | 29–55 | 56–85 | >85 |
Liability sensitivity: Very High (public accountability) | Regulatory pressure: High | Litigation exposure: High (constitutional)
| Category | Debt Free (<) | Manageable | Dangerous | Critical (>) |
|---|---|---|---|---|
| D1: Bias | <5 | 5–10 | 11–16 | >16 |
| D2: Transparency | <5 | 5–10 | 11–16 | >16 |
| D3: Governance | <6 | 6–11 | 12–17 | >17 |
| D4: Privacy | <6 | 6–11 | 12–17 | >17 |
| D5: Accountability | <5 | 5–10 | 11–16 | >16 |
| Total | <27 | 27–52 | 53–82 | >82 |
Liability sensitivity: High (brand-driven) | Regulatory pressure: Medium-High | Litigation exposure: High (class action)
| Category | Debt Free (<) | Manageable | Dangerous | Critical (>) |
|---|---|---|---|---|
| D1: Bias | <7 | 7–13 | 14–19 | >19 |
| D2: Transparency | <7 | 7–13 | 14–19 | >19 |
| D3: Governance | <7 | 7–13 | 14–19 | >19 |
| D4: Privacy | <6 | 6–12 | 13–18 | >18 |
| D5: Accountability | <7 | 7–13 | 14–19 | >19 |
| Total | <34 | 34–64 | 65–94 | >94 |
The Liability Ledger is part of an integrated framework ecosystem. Your Liability Ledger score directly undermines your Trust Premium score — they are inversely correlated. High ethical liability destroys trust premium.
Assessment The Trust Premium Assessment — Quantifying the business value of trusted AI (15 dimensions, 0–75 score) Governance Minimum Viable Governance (MVG) — The smallest governance structure that prevents the largest AI failures Assessment The PRIME Assessment — Measuring organizational AI readiness across five dimensions Strategy The 5-Pillar AI Strategy — Building enterprise AI strategy on five foundational pillars Playbook The AI Governance Playbook — Operationalizing responsible AI from policy to practiceYou do not need the full Liability Ledger assessment to start. This 90-minute quick audit produces a directional score and identifies your highest-compound-rate debts. Assemble 3–5 people: one AI/data science lead, one legal/compliance representative, one business owner, and optionally a risk officer and HR representative. Based on the methodology from Measuring Ethical Debt: A Practical Scoring Method.
| AI System Name | Owner | Production Date | Last Reviewed | Type (Internal / Vendor / Shadow) |
|---|---|---|---|---|
| | ||||
| | ||||
| | ||||
| | ||||
| | ||||
| |
| AI System | D1 Bias 1–5 |
D2 Transp. 1–5 |
D3 Gov. 1–5 |
D4 Privacy 1–5 |
D5 Account. 1–5 |
Total 5–25 |
|---|---|---|---|---|---|---|
| AI System | Category | Raw Score | Months in Prod. | Compound Rate | Periods (n) | Compounded Score Score × Rate^n |
|---|---|---|---|---|---|---|
| Rank | AI System | Category | Compounded Score | Why This Is Priority |
|---|---|---|---|---|
| 1 | | |||
| 2 | | |||
| 3 | |
| Rank | AI System + Category | Named Owner | Deep Audit By | Remediation Plan By | Executive Sponsor |
|---|---|---|---|---|---|
| 1 | |||||
| 2 | |||||
| 3 |
Match the tool to the score. Critical-debt organizations (score 4–5) should start with free tools to establish baseline visibility. Well-managed organizations (score 1–2) should invest in enterprise platforms for continuous automated monitoring. Based on the tools landscape from Measuring Ethical Debt.
| Category | Free / Open-Source Tools | Enterprise Tools |
|---|---|---|
| D1 Bias 2.0x compound |
IBM AI Fairness 360 (AIF360) — 70+ fairness metrics, 10 mitigation algorithms Fairlearn (Microsoft) — additional mitigation algorithms |
Fiddler AI — real-time bias detection, compliance dashboards Truera — model intelligence platform Credo AI — AI governance and compliance |
| D2 Transparency 1.3x compound |
SHAP — feature importance / Shapley values LIME — local interpretable model explanations Model Cards — standardized documentation |
Fiddler Explainability — production explainability Arthur AI — model monitoring + explainability MLflow / Weights & Biases — model registry |
| D3 Governance 1.5x compound |
MVG Framework (askajay.ai) — 90-day implementation NIST AI RMF — compliance crosswalk Internal audit cadence — quarterly minimum |
ModelOp — enterprise model management Holistic AI — AI governance platform Credo AI — governance + risk management |
| D4 Privacy 1.8x compound |
DPIA templates — Data Protection Impact Assessment Data classification frameworks ICO AI guidance — AI-specific data protection |
OneTrust — privacy management platform BigID — data intelligence + privacy Collibra — data governance + cataloging |
| D5 Accountability 1.5x compound |
RACI templates — ownership clarity Decision log templates — attribution EDPB AI auditing checklist — regulatory baseline |
ServiceNow GRC — governance, risk, compliance Archer (RSA) — integrated risk management Diligent — board-level governance platform |
Quarterly is the minimum cadence for formal AI governance audits. At a 2.0x compound rate, Bias Debt doubles in six months. A quarterly audit catches the debt at 1.4x — before it doubles. Track your scores here to measure debt reduction over time and demonstrate governance progress to the board.
| Quarter | D1 Bias /25 |
D2 Transp. /25 |
D3 Gov. /25 |
D4 Privacy /25 |
D5 Account. /25 |
Total /125 |
Change vs. prior |
Auditor |
|---|---|---|---|---|---|---|---|---|
| Q1 | ||||||||
| Q2 | ||||||||
| Q3 | ||||||||
| Q4 |
The same intellectual honesty that the canonical concept article applies to the framework applies to this worksheet. Four objections deserve direct engagement before any score is presented externally.
1. “Compound rates are not empirically derived from a closed-form model.” Acknowledged. The 2.0x / 1.8x / 1.5x / 1.3x rates encode the directional convergence of multiple independent data streams — EEOC enforcement velocity, GDPR fine compounding, Workday-class third-party liability rulings, Apple Card-style transparency precedents. They are not the output of a single regression. The rates communicate relative urgency, not insurance-grade probabilities.
2. “The data mixes verified and partially verified statistics.” Acknowledged. The canonical articles distinguish between fully verified figures (Meta $1.4B, EEOC $365K, Nature 91% degradation, IBM $670K shadow-AI premium) and partially verified ones (estimated error-rate increases, settlement values denominated in equity). Where a figure has caveats, they are noted in the source articles. The directional argument does not depend on any single statistic.
3. “Inverted scoring is unfamiliar and creates user-error risk.” Acknowledged. Most maturity assessments score higher = better. This one scores higher = worse debt. The Inverted Scoring callout near the top of the worksheet exists for this reason. Score reviewers should explicitly confirm orientation when comparing Liability Ledger scores against other framework scores in the same governance dashboard.
4. “The framework cannot capture jurisdiction-specific nuance.” Acknowledged. EU AI Act, US state AI laws (Colorado, NYC AEDT, California), UAE PDPL Article 18, and sector-specific rules (HIPAA, GLBA, FCRA) each impose different obligations on the same operational behavior. The worksheet provides a portable diagnostic. Jurisdiction-specific compliance requires legal counsel and the relevant regional companion guides linked in the Evidence Base.
Definitions used throughout this worksheet. These align with the canonical Liability Ledger articles.
Liability Ledger. A diagnostic framework for inventorying, scoring, and reducing the hidden liability accumulating in an AI portfolio. Five debt categories (D1–D5), five sub-dimensions each, scored 1–5 (inverted). Total range: 25 (no debt) to 125 (critical debt across the board).
Ethical Liability Score. The sum of all 25 sub-dimension scores. Lower is better. Maturity bands: 25–50 Well-Managed, 51–75 Monitored, 76–100 Material Debt, 101–125 Critical Debt.
Compound Rate (Interest Rate). The multiplier by which the cost of remediating a debt category rises every six months it is deferred. Bias 2.0x, Privacy 1.8x, Governance 1.5x, Accountability 1.5x, Transparency 1.3x. Directional, not actuarial.
D1 Bias Debt. Discriminatory outcomes in AI systems — hiring, lending, housing, healthcare. Fastest-compounding category, driven by EEOC enforcement and Workday-class third-party liability.
D2 Transparency Debt. Unexplainable models, undisclosed AI use, missing documentation. Compounding accelerates as EU AI Act high-risk-system obligations take effect (August 2026).
D3 Governance Debt. Shadow AI, missing inventories, no risk tiers, no designated owners. The multiplier on every other category — you cannot pay down debts you cannot see.
D4 Privacy Debt. Consent gaps, biometric data practices, cross-border transfer violations. Second-fastest compound rate; drives the largest historical fines (GDPR cumulative >€5.88B through 2024).
D5 Accountability Debt. Missing human oversight, no escalation paths, unclear vendor liability. Hides in vendor contracts; Workday ruling shows contractual liability caps may not protect against third-party discrimination claims.
Priority Paydown Plan. A 90-day sprint that addresses the highest-interest-rate debt first. Pay down Bias before Transparency, Privacy before Accountability — not because Transparency does not matter, but because compounding determines remediation cost.
The Workday Precedent. Federal court ruling that an AI vendor (Workday) could be liable for discrimination claims brought by job applicants who never signed a contract with the vendor — expanding third-party AI liability beyond contractual privity. Drives the Accountability Debt compound rate.
This worksheet is the operational layer of a published, sourced framework. The methodology, compound-rate derivation, enforcement-action evidence, and counter-evidence discussion live in the canonical articles below.
Canonical articles:
Companion frameworks:
Key external sources: EU AI Act Article 99 (penalties) · IBM 2024 Cost of a Data Breach Report · EDPB / GDPR enforcement tracker · EEOC AI & algorithmic fairness enforcement