Glossary

Cortexa AI Glossary · Trust, fakes, and safety

Can we see inside an AI (explainability)?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Even the builders can't always say why an answer came out. Researchers are starting to look inside.


The dog in the cats folder

Your photo app files a picture of your dog under "cats." You can't ask it why, and the people who built it often can't point to the exact reason either. Inside an artificial intelligence (AI) model are billions of numbers, tuned during training, and none of them is labeled "because." That's why people call these systems black boxes. Researchers are working on opening the box, and they've made a real start.3

Two ways to look

Researchers use two words for this work, and they mean slightly different things.

  • Explainability: methods that point to what most affected an answer, such as which words in a question or which parts of a picture mattered most.
  • Interpretability: looking at the inside of the model directly, to see what its numbers are doing.

A heat map laid over your dog photo is explainability. It might show that the app was staring at the sofa behind the dog.1

A map of concepts

In May 2024, the AI company Anthropic reported one of the most detailed looks inside a large language model so far. Its researchers studied a middle layer of Claude 3 Sonnet, one of its own chatbot models. They found patterns of activity, which they call features, that line up with ideas a person would recognize. Millions of them. One lit up for the Golden Gate Bridge, whether the bridge came up in English, in Japanese or Russian, or in a photo.2

Turning the dial

Finding a pattern is one thing. Showing that it does something is harder. So the researchers turned that bridge feature up, the way you might turn up a dial. The model started bringing the bridge into answers where it didn't belong. Asked about its own form, it said it was the Golden Gate Bridge. It sounds like a joke, but changing the feature changed what the model said, so the feature is part of how the model works.2

Only part of the map

The researchers were careful about limits. They wrote that the features they found are a small subset of all the concepts the model learned, and that finding every one with their method would take more computing than training the model did. A 2025 follow-up traced how the model handles single questions. Even on short, simple prompts, it captured only a fraction of the work, and reading each one took a person a few hours. So we can see some of the wiring. Nobody can yet trace the whole reason behind a big model's answer.24

Ask it why

That follow-up found something useful for everyday use. The researchers watched the model add 36 and 59. Inside, it ran two paths at once: one made a rough estimate, and the other worked out the last digit. Then they asked how it got the answer. It described carrying the one, the method taught in school. Its account sounded right, and it didn't match what had happened inside. A model's explanation of its own answer is more text it wrote, and you can check it like any other claim.4

Why people keep looking

Looking inside could help researchers catch a model leaning on a bad shortcut, find hidden bias, or confirm that a safety fix changed what it was meant to change. Anthropic says that's why it does this work: to make models safer. Next time an app sorts something oddly, what do you think it was looking at?12

Works cited

  1. IBM, "What is AI interpretability?" (checked )
  2. Anthropic, "Mapping the mind of a large language model" (2024) (checked )
  3. IBM, "What is black box AI?" (checked )
  4. Anthropic, "Tracing the thoughts of a large language model" (2025) (checked )