← Back to insights Measurement

AI fluency - are you using AI effectively and safely in your organisation?

Usage numbers tell you how much AI is being used but they do not tell you if anyone is using it well and improving. The sanctioned route into your organisation's Claude activity - the Compliance API - cannot answer that question, and it's important to understand why before you rely upon it on its own.

Insight · Measurement · Rubytech

The author, Joel Smalley, has completed Anthropic's AI Fluency: Framework & Foundations, Teaching the AI Fluency Framework and AI Fluency for Small Businesses (Anthropic and PayPal).

Businesses measure AI adoption by the number of seats taken up, messages sent, share of staff active in the last week, etc. The number goes up, it is treated as progress. However, it is not progress, it is attendance. Adoption and skill are different things, and skill - or fluency to adopt the industry term - is the one the business is paying for. The question a leader actually needs an answer to is whether their staff are getting better at using AI productively, and indeed, whether that can be measured at all.

What the official route can see

Anthropic's Compliance API gives you two things. An audit feed of events, which is who did what and when, at the level of the account. And, on Enterprise, the conversation content itself: the chat text, together with the files and artifacts that came out of it.

That is a real requirement and it is the right foundation for the job it was built for - data compliance, legal discovery, retention policy, and so forth.

It does not, however, give you the tool-use trace. In other words, it does not tell you which tools were called, with what inputs, or what came back. And tool use is the biggest risk that comes from AI deployment. It deserves special attention. It also does not carry the reasoning that led from the question to the answer. You get the final outputs but not the working - the thought process of both the AI and the human operator. At the admin level, you can assess whether the answer was reasonable given the prompts but you cannot see, and therefore assess, whether the person knew what they were doing - either in terms of domain knowledge or AI proficiency.

So a fluency assessment built on the Compliance API is deficient by design.

What each source can see
 
Compliance API
Local transcript
Reflection log
Usage and events
Full
Full
None
Conversation text
YesEnterprise
Full
None
Files and artifacts
YesEnterprise
Full
None
Tool-use trace
None
Full
None
Reasoning and iteration
None
Full
None
What they did with the output
None
None
Only source
The compliance route carries the outputs and stops. The trace carries the working and stops at the end of the session. The last row is the one nobody logs.

The fluency information is in the local transcript

The material that answers the question is already being recorded on the user's local device. A Claude Code session keeps its own log, a plain text file with one record per line (the “JSONL”), and that log holds the full trace: every tool call and the input it was given; every result that came back; the corrections, the dead ends, the point where the person intervened, challenged the AI's assumptions or even reconsidered their own.

That is where the behaviours live - the “discernment” part of the 4D framework that Anthropic teach as part of their fluency courses - whether someone gave enough context to make the task well posed, whether they questioned a result that looked correct because it was confidently articulated, whether they narrowed the work into steps or handed over the whole thing and hoped. None of that is visible in the finished output but all of it is visible in the session transcript.

Those logs sit on your staff's machines and you can collect them. And you should!

Consent as part of the governance process

Collecting session logs is a governance matter before it is a technical one. Transparency is an integral part of the “Diligence” part of the 4D framework, not least because consent and transparency are part of diligent governance independently of AI. Staff should be informed about what is collected, why, who reads it, how long it is kept, and that the purpose is their development rather than their supervision. It is part of the learning process towards AI fluency as much as it is the responsibility of the administrator to keep good records and monitor such fluency in the interest of the business.

The limitations of the session transcript

The transcript shows what happened inside the session. However, this is still not entirely sufficient for the assessment of AI fluency. It does not show what the person did next: whether they checked the figures against the source; whether they told the client the draft was AI-written; whether they read the output properly or simply forwarded it verbatim; whether they were right to trust it. None of this is in the log but it is where most of the judgement sits, a critical part of the fluency assessment.

Anthropic's own research says the same thing about its own platform. The AI Fluency Index uses the 4D framework developed by Rick Dakan and Joseph Feller, which defines 24 behaviours of effective work with AI. Eleven of them are directly observable when someone uses Claude. The other thirteen, including being straight about AI's role in a piece of work and thinking about the consequences of passing the output on, happen away from the interface, and Anthropic notes that these are arguably the most consequential of the set.

Roughly half the thing you want to measure is not on the platform at all.

Figure 2 · The 24 fluency behaviours, by dimension
In the transcript Reflection only
Delegation2 of 7 in the transcriptDecide what to do with AI vs. yourself.
Clarifies the goal Consults on approach Spots a poor fit Knows AI's limits Stays involved Picks the right tool Rebalances the work
Description6 of 8 in the transcriptCommunicate clearly with AI.
Defines the audience Specifies the format Communicates tone Builds iteratively Provides examples Sets the interaction Breaks into steps Sets boundaries
Discernment3 of 5 in the transcriptEvaluate what AI gives you.
Checks facts Questions the reasoning Spots missing context Tracks progress Notices style mismatch
Diligence0 of 4 in the transcriptUse AI responsibly and accountably.
Guards what's shared Straight about AI's role Owns the output Weighs the consequences
11 land in the transcript. 13 need the reflection.
The behaviours and the dimension definitions are Anthropic's own. Delegation and Description carry most of what a session can show. Diligence carries none of it, and Anthropic calls those four arguably the most consequential of the set.

What is necessary for complete AI fluency assessment

A complete measure looks through three lenses: the outcome, the process and the reflection.

The outcome is the person's final work set against the AI's last draft. What they changed, what they cut, what they checked and what they let stand is the objective record of whether their oversight added value. It is the one lens that asks nothing of the person beyond the two documents.

The process is the transcript. It gives you what happened: the tools, the steps, the iteration, the corrections, the path from question to result.

The reflection is a short log kept by the person and it gives you the rest. A running journal of a few lines per piece of work, or a brief reflective session at the end of every AI collaboration, answering three questions:

  1. What I asked for, what I did with what came back;
  2. What I checked and corrected;
  3. What I would do differently next time and how I have recorded this to be a durable improvement.

It is the person's own account, which makes it weaker evidence than the trace and the only evidence there is for the half the trace cannot reach. That is why it sits alongside the other two lenses rather than on its own.

A fluency measure is three things - the outcome against the AI's last draft, the trace of the session, and the person's account of what they did with it. No one of them is a measure on its own, and a usage report is none of them.

However, none of this is readily available. The information retrieval, the consent frame, the reflection process, the outcome comparison and the compilation of the three sources all have to be assembled.

Sources
  1. Anthropic Compliance API: scope of the audit event feed and, on Enterprise plans, conversation content including files and artifacts. Anthropic developer documentation, platform.claude.com/docs, Compliance API pages, read July 2026.
  2. Anthropic Education Report: The AI Fluency Index (Kristen Swanson, Drew Bent, Zoe Ludwig, Rick Dakan and Joe Feller, 16 February 2026): the 4D AI Fluency Framework's 24 behaviours, of which 11 are directly observable on Claude.ai and Claude Code and 13 occur away from the interface. anthropic.com/research/AI-fluency-index
  3. The 4 Ds of AI Fluency, behavioural indicators: the reference list of the fluency behaviours cited in the Index. Anthropic Tutorials, claude.com/resources/tutorials.

Assembling it is what the training covers. A decision-maker or an administrator learns to work out a policy for measuring fluency and governing AI use that suits the size of their firm, and to put it in place themselves. Small firms need that as much as large ones, and generally have less to assemble.

If you want to know where you stand before any of that, the free capability and governance audit is fifteen questions and gives you your scores on screen.

Start a conversation More insights