The Glass Box Paradox: When AI Can’t Explain Itself (Inside my AI Law & Policy Class #10)
AI Transparency, When Seeing Isn’t Understanding
8:55 a.m., Monday. Welcome back to Professor Farahany’s AI Law & Policy Class.
You’re 26, just graduated, and are $80K in debt. An AI rejects your job application in 0.3 seconds. Another AI denies your apartment rental. A third flags you as high risk for health insurance.
None will tell you why.
Today we’re examining AI transparency—the solution everyone claims to want but nobody agrees how to implement. Where billion-dollar trade secrets collide with public accountability. Where “explainable AI” meets “that’s proprietary.” Where laws require disclosure of things that may be technically impossible to disclose.
Consider this your official seat in my class—except you get to keep your job and skip the debt. Every Monday and Wednesday, you asynchronously attend my AI Law & Policy class alongside my Duke Law students. They’re taking notes. You should be, too. And remember that live class is 85 minutes long. Take your time working through this material.
Just joining us? Go back and start with Class 1 (What is AI?), Class 2 (How AI Actually Works), Class 3 (Open vs. Closed AI Models), Class 4 (The Control Paradox), Class 5(The $1.5 Billion Question), Class 6 (Training Data, Discovery Wars, and Who Gets Paid), Class 7 (Why China Quit US Chips), and Class 8 (Your Electricity Bill and Compute), and Class #9 (AI Bias). At the very least, do classes 1, 2, 5, and 6. That foundation will help you understand why today’s topic matters so deeply.
Let’s start with something that probably happened to you this week. You applied for something online—maybe a 0% interest credit card. Within seconds, you got rejected. No explanation. Just “unfortunately, we cannot move forward at this time.”
Now I’m going to show you three “transparent” explanations for why that happened. Tell me which one actually helps you:
Explanation A: The system used a deep neural network with 175 billion parameters trained on 45TB of text data using transformer architecture.
(Translation: We used fancy math you wouldn’t understand.)
Explanation B: The AI looks for patterns similar to ones it’s seen before and makes a guess based on those patterns.
(Translation: The computer noticed something about you that reminded it of other people we rejected.)
Explanation C: Your credit application was denied because your debt-to-income ratio of 47% exceeds our threshold of 43%.
(Translation: You spend too much of your paycheck on debt payments. Pay off $200/month and reapply.)
In my live class today with your Duke Law counterparts, they all chose C. You probably chose C, too, right? Finally, something actionable!
But here’s the gut punch that the in-person and you students now get. That’s not how most AI systems actually work. They can’t give you Explanation C because they don’t “know” why they made the decision in those terms.
You know how you sometimes get a “vibe” about someone but can’t explain why? The AI made millions of tiny calculations that together produced a “no”—but those calculations don’t translate to human concepts like “debt-to-income ratio.” The system might have noticed you shop at stores that correlate with default risk, combined with your zip code’s unemployment rate, mixed with the timing of your credit card payments, plus 9,997 other factors.
So that explanation “C” is probably just made up after the fact by humans trying to approximate what the AI might have been thinking.
In the in-person class, we talked about how humans rationalize their decision-making, too. I asked my students, for example, to explain why they chose Duke Law. One student said it was because of the employment statistics. Another one said it was because of the trees in May during their visit. The truth? Those were probably one factor among many that drove their decisions and even they don’t know why they made the decision that they did.
Which brings us to a central question for today. If AI can’t explain itself at all, let alone in terms that humans understand, can it ever truly be transparent?
Part I: Technical Transparency—What Can We Actually See?
In May 2024, Anthropic—the company behind Claude—announced something wild. They’d finally looked inside their AI’s “brain” and found it knows things about you that you might not even know about yourself.
They found what they call “features”—patterns that light up when the AI thinks about specific concepts. Some are obvious. They found a feature for “Golden Gate Bridge.” Cool, the AI knows about bridges.
But then it got more surprising. They found a feature for “deception.” Another for “writing insecure code.” Features that detect emotional manipulation, financial distress, relationship problems.
This sounds like breakthrough transparency, right? We can finally see what the AI is thinking!
But there’s a catch, which is that they only mapped 1% of Claude’s features. The other 99%? Complete mystery.
Think of it this way:
That job application AI that rejected you? 99% of its decision-making is unknowable
The AI that sets your insurance rates? 99% black box
The algorithm determining your financial aid package? 99% mystery meat
Think about exploring your apartment with your phone’s flashlight during a power outage, with motorized blackout shades stuck in the closed position. You use your phone to light up the bathroom. Awesome! You know where the toilet is. But your kitchen, bedroom, that weird closet that was locked when you moved in? Total darkness. And something’s making noise in there.
The Three Windows Exercise
Let’s make this concrete with something that might have happened to you or someone you know.
Sarah, 28, software engineer, makes $95,000 a year. Perfect payment history. Just got denied for a mortgage. The bank offers three types of “transparency”:
Window 1: Process Transparency. We collected your credit data, ran it through our AI system, and it produced a decision.
Sarah learns an AI was involved, and the processing stream, but nothing else.
Window 2: Technical Transparency. The neural network processed 10,000 data points through 50 billion parameters using matrix multiplication and gradient descent optimization.”
Even Sarah, who codes for a living, has no idea what to do with this information.
Window 3: Outcome Transparency. You were denied because of: late payments (40%), high credit utilization (30%), short credit history (30%).
Finally! Actionable information! Except... Sarah has never made a late payment. She uses only 10% of her available credit on her credit cards. Her credit history dates back 10 years. These percentages were made up after the fact to approximate what the AI might have been considering. The AI doesn’t actually think “40% late payments,” it noticed that Sarah shops at the same stores, lives in the same zip code, and has similar browsing patterns to other people who defaulted on loans.
This is what Ben Chester Cheong calls the “explainability gap”—it’s not just hard to explain AI decisions, it may be theoretically impossible because AI doesn’t “decide” the way humans do.
In the live class, at this point, the student start to see where this is going. Just like every other class this semester, what we mean when we say transparency can be a lot of different things. And we better make sure our laws are actually targeting the window(s) we care about.
Part II: Legal Transparency
If we accept that AI can’t explain itself, should it be making decisions about people’s lives? Keep this in mind as we look at what was introduced into the U.S. Congress tried to require (as one of many examples of transparency laws that we’ll look at this semester):
H.R. 7532: Congress Tries to Mandate the Impossible
While Anthropic was discovering they could only see 1% of Claude’s brain, Representative Comer was writing H.R. 7532—a law requiring federal agencies to explain their AI decisions.
Remember our Three Windows? The bill tries for all three:
Window 1 (Process): Agencies must say they used AI ✓
Window 2 (Technical): Must describe training data and “logic” ✓
Window 3 (Outcome): Must notify affected individuals and allow appeals ✓
But the core requirement, of making AI “sufficiently explainable and understandable,” came with a crucial caveat of “to the extent practicable.”
And overall, the bill focused more on institutional transparency (what systems exist) than decisional transparency (why you were denied). Because it’s a lot easier to list your AI tools than explain how they think.
The bill does include enforcement—Inspector General reviews, public posting requirements, appeal rights. But when the core requirement is “explain the unexplainable, unless you can’t, the enforcement mechanisms are checking whether agencies filed paperwork.
H.R. 7532 perfectly illustrates our transparency paradox. It mandates:
Notification when AI affects you (doable)
Technical documentation (possible but doesn’t help you figure out why you were denied a loan)
Meaningful explanation (not really possible with current neural networks)
The downside is that we end up in a kind of “transparency theater,” where agencies produce massive documentation that technically complies but helps nobody. Citizens get notices saying “AI was involved in your denial” with no actionable information. The appeal process? You argue to humans who also don’t understand how the AI decided.
(The bill stalled in committee. Not because of industry opposition (they could work with “practicable”), but because even supporters realized it was requiring something technically impossible while creating massive bureaucratic burden.)
Part III: Political Economy—The Global “Solution”
February 2025. While Congress was burying H.R. 7532, the G7 launched the Hiroshima AI Process Reporting Framework. This voluntary reporting framework is meant to encourage transparency and accountability among organizations developing advanced AI systems and, according to the OECD, to “facilitate transparency and comparability” of risk mitigation measures.
The Seven Pillars of the Hiroshima AI Process Reporting Framework include:
1. Risk identification
2. Risk management
3. Transparency (yes, transparency about transparency)
4. Governance
5. Content provenance
6. Research for AI safety
7. Advancing global interests
This looks promising! But now we are in a voluntary commitment world, where companies choose whether to participate, if so, what they are willing to disclose, when they will stop disclosing information, and how to interpret the requirements. (It’s like having speed limits that are suggestions.)
Curious what reporting looks like? Check out Anthropic’s submission, where they cross-reference their “Transparency hub.”
The OECD Report Card: Grading on a Curve
The OECD analyzed how companies are actually using Hiroshima. They looked at Anthropic, Google, Microsoft, OpenAI.
The Headline from a few days ago: “Companies are taking significant steps toward transparency!”
The Reality (on p. 7 of the OECD report): “Technical provenance tools such as watermarking are still limited in deployment, face limits, and are mainly used by large technology firms.”
Meaning the biggest companies are disclosing some measures but robust risk mitigation strategies remain sparse.
And as you dig more deeply into the report, what you see is:
What Companies Actually Report:
“We have a process” ✓
“We think about safety” ✓
“We’ve identified some risks” ✓
What Kind of Transparency Would Help (but we don’t see):
WHICH risks they found ✗
HOW OFTEN things go wrong ✗
WHO got harmed ✗
WHAT they’re doing about it ✗
Part IV: Real-World Transparency
NYC Local Law 144: When Transparency Makes Things Worse
Remember that time you didn’t get a job and wondered if it was discrimination? Well, New York City decided to solve that with the radical transparency approach we discussed in class #9. Now companies must post their AI bias audits publicly.
Here’s an actual posting from a major tech company (name redacted but you’d recognize them):
AI Hiring Tool Audit Results (2024):
White candidates selected: 8.2%
Asian candidates selected: 7.9%
Black candidates selected: 6.4%
Hispanic candidates selected: 6.5%
Women selected: 6.8%
Men selected: 8.1%
The legal threshold for discrimination is 80% of the highest rate (that’s 0.8 × 8.2% = 6.56% in this case).
In the live class, one of the student asked … “if I am a cross between two of those categories (woman and black), what does that tell me about my chances of applying?
(Unclear).
And then we talked about … if you are an applicant who was rejected, and you see this report, does it help you? Hurt you?
Imagine you’re a Black female software engineer. You see this posting. The company is literally telling you their AI is calibrated to discriminate against you as much as legally possible.
This is transparency! They’re being honest! But is it somehow worse than not knowing? Especially if the law doesn’t give you a right of action, and it doesn’t satisfy other civil rights statutes for a claim?
What’s particularly hard is that transparency can now also drive companies to optimize their AI to hit these numbers precisely. Not to be fair, not to find the best candidates, but to discriminate exactly at 79.99% to avoid lawsuits while still filtering out as many protected class members as possible.
Houston Teachers: When Full Transparency Costs $2 Million to Understand
Here’s a story that will make you appreciate why your professor uses a simple grading rubric.
Houston schools used a “value added” algorithm to evaluate teachers. It was so complex that when teachers sued for transparency and won, here’s what they got:
The “Transparent” Algorithm Package:
900 pages of mathematical formulas
10,000 lines of code
Statistical models requiring a PhD to interpret
Documentation written entirely in technical jargon
The Price Tag to Understand It:
Expert witness to translate: $500/hour
Statistical analysis: $300,000
Legal interpretation: $200,000
Total cost to understand their own evaluation: $2 million
Teachers’ Budget: $50,000
It’s like asking why you failed a test and being handed the source code to Microsoft Word. Technically it’s all there, but practically it’s useless.
Part V: Synthesis—Making Sense of the Mess
Picture this as a three-layer cake of dysfunction:
Bottom Layer: Technical Reality
We can only see 1% of how AI works
That 1% doesn’t translate to human understanding
True transparency may be impossible
Middle Layer: Legal Response
H.R. 7532: Mandated impossible transparency, failed
Laws require explaining the unexplainable
Every requirement has an escape hatch
Top Layer: Political Layer
Hiroshima: Voluntary standards (with significant geopolitical limits on how much any superpower wants to disclose)
Companies control disclosure
Transparency theater protects power
The Five Levels of Meaningful Transparency
Which is why, perhaps instead of asking “is the AI system transparent?” we might want to ask, when denied a loan or another adverse decision is made, “what can a person DO with this information at each level of transparency?” (For a nice discussion of this concept, check out Mike Hsu’s post “AI Actionability Over Interpretability”).
Think of this as a pyramid. Most AI transparency stops at the bottom level, but real accountability requires climbing toward the top.
The Five Levels of Meaningful Transparency
Level 1: Notification. “AI was used” to make the decision.
You get a tiny disclaimer: “This decision involved automated systems.”
Example: Your insurance denial says, “processed using automation.” Cool. Super helpful.
What you can do: Nothing. It’s like being told “we decided.”
Who does this: Everyone. Bare minimum legal compliance.
Level 2: Explanation. “Here’s roughly what happened”
You get vague categories: “We considered credit, income, and other factors.”
Example: “Your risk score exceeded our threshold based on financial factors.”
What you can do: Guess what went wrong. Was it income? Credit? The mysterious other?
Who does this: Some lenders, NYC’s audit law, H.R. 7532’s goal.
Level 3: Interrogation. “What if I changed X?”
Now you can test scenarios: “Would $10K more income change the decision?”
Example: Adding a co-signer with 720+ credit score = approval
What you can do: Actually fix the problem. Spot errors (My income isn’t $0!)
Who does this: Almost nobody. Would reveal the secret sauce.
Level 4: Contestation. “I dispute this”
You can appeal with evidence. A human reviews and can override.
Example: “You say I’m unemployed. Here are five years of W-2s.”
What you can do: Fix mistakes. Add context AI missed. Get human judgment.
Who does this: Europe’s GDPR (theoretically), some banks (grudgingly).
Level 5: System Change. “Your dispute fixed the model”
Your appeal doesn’t just help you, it fixes the AI for everyone.
Example: Your case revealed bias against self-employed applicants. We’re retraining and reviewing past decisions.
What you can do: Create systemic change. Help others. Force real accountability.
Who does this: Nobody. Admitting systematic discrimination = massive liability.
Each level gives you more POWER, not just information. Most companies camp out at Level 1-2, giving you knowledge without agency. Real transparency requires Level 4 minimum, the ability to act on what you learn.
The Uncomfortable Truth(s) For Today
Perfect transparency is impossible (we can’t explain what we don’t understand)
Partial transparency might be worse (like NYC’s discrimination ratios)
But zero transparency is unacceptable (democracy requires accountability)
And all of this helps us see what comes next … which is accountability and liability. What transparency might enable.
Your Homework
Share this with one person you think should know more about AI transparency
Find an AI decision that affects you (credit score, job application, insurance rate, anything). Try to get an explanation for how the decision was made. Share your findings in the comments. Let’s map the real state of AI transparency together.
Class dismissed. But remember. Every time an AI makes a decision about your life, it’s doing math it can’t explain, following logic it doesn’t understand, creating outcomes nobody can fully trace. The question isn’t whether that’s transparent—it’s whether that’s acceptable.
Class dismissed! But we are not done yet. There’s still a LOT to cover in AI Law and Policy.
And if you’re curious … this is what the board look like from live class. You get better visuals here.
The entire class lecture is above, but for those of you who find this work valuable and want to support me (thank you!), or who want to go deeper in the class, the class readings, video assignments, and virtual chat-based office-hours details are below.







