The AI Agent Factory

6.7 Review work you did not watch

Status
stable
7 min read
Owner
Panaversity
Approved
Panaversity ·

In everyday life. Someone looking after your home while you are away texts "all fine." You check the photos, the mail and whether the plants were watered.

Delegated work happens while you are away. You review a result and a record, not a process you saw. The AI Fluency Framework, by Rick Dakan and Joseph Feller with Anthropic, calls this skill discernment.1 Discernment means judging AI's work, and the framework splits it three ways. Product discernment judges what came back. Process discernment judges how the AI got there. Performance discernment judges how it behaved while working.2 The checks in 6.2 to 6.6 of this chapter are product discernment. You do the other two by reading the record.

A finished run is not a correct run. A status that says "completed, no errors" means the task ended without a technical failure. It says nothing about whether the work is right. Read the evidence your contract asked for, not the status.

Read the task record, the worker's record of the run, for three things.

  1. What it used. Which files it read and which messages it received. Did it read the governed policy, the approved version in use? Did it get the decisions that Maria, the office manager, made?
  2. What it did. Which files it wrote and which messages it sent. Did it change anything it was told not to change?
  3. What it could not have done. Compare the claims in the output with the tools in the record. Brightline's AP Worker handles the bills the company owes. Its output for Friday's payment run said "Bank change verified with Tri-County Freight." Tri-County, a supplier, had asked to be paid into a new bank account. The worker's record shows two tools, the only actions it could take: read files and write files. It also shows Maria saying "I will call," and on Thursday she told Dave, the controller, that she had not. The worker could not have made the call. No person had made it, and nothing records a call. So the claim is fabricated: the worker invented it.

Check the record itself, too. If the worker writes its own summary of what it did, compare that summary with any log the AI product keeps, or with the source system, such as the accounting system, when you can.

You can ask a second worker to review the first. It is good at finding possible problems fast. It is not the review. It can miss what you would find and invent problems that are not there. It can also mark a correct abstention, an honest "the source does not say," as a gap. Use it to search more widely, then check each finding yourself. The person who signs is responsible for the result.

The title reads "Check the claim against the record," and the line below it reads "A finished run is not a correct run." Two panels are joined by a broken chain link. On the left, the output claims "Bank change verified with Tri-County Freight," and its proposed action is to pay to the new account. On the right, the evidence shows three things. Tools available: read files and write files, no phone or email. Maria's instruction on Monday: "Hold both Tri-County invoices, and I will call." Maria's confirmation on Thursday: she has not made the call. Both panels lead to a red box: the verification claim is fabricated. The worker could not call, and Maria confirms she had not called. Below, under read every task record for three things, three boxes. One, what it used: files read and messages received. Did it use the governed policy and Maria's decisions? Two, what it did: files written and messages sent. Did it stay within its instructions? Three, in gold, what it could not have done: compare claimed actions with available tools, and seek independent confirmation where needed. A footer reads: "Completed, no errors" reports execution status, and does not establish correctness. Cross-check worker-written summaries against product logs or source systems when available.

Figure 6.6. Check the output's claims against the record.

Check yourself

Question 1 / 8 · blueprint

0 answered

A task record says 'completed, no errors.' What does it tell you?

  1. AI Fluency: A Crash Course, Panaversity, first edition.↑
  2. AI Fluency: Key Terms, Rick Dakan, Joseph Feller and Anthropic.↑

On this page