I Want to See My Agents Work
I do not need to watch every token an agent produces. I do need to be able to see that my agents are working, which agents are doing which parts of the job, and where they are in the process. When the run is over, I need to be able to inspect how they arrived at what they produced.
That is a different expectation from the one most AI products set. You give a system a request, it disappears into a black box, and it returns with output. The output may be good, maybe even good enough to use. But if the work was consequential, the person who approved it is still responsible for it. Responsibility without visibility is incomplete.
In The Cheaper Half of Oversight, I wrote about reviewing the plan before agents begin. That is where you decide whether the system is doing the right work. This is the next part of the same problem: once the work begins, the agents should be visible workers, not black boxes that simply return output.
What I Need to See During a Run
In a four-agent Compound Console run, the first useful fact is that four distinct specialists are active, each with a defined assignment in the approved plan. One may be researching the market while another is analyzing competitors. A third may be turning their findings into a recommendation. The run graph and the agent cards show which work is underway, which work is complete, and which work is waiting on another agent, so you can see how much of the approved sequence is still outstanding.
That is basic operational awareness. A manager does not need to read every document a team member writes in order to know whether the work is moving. They need to know who owns the work, where it sits, and whether the operation is progressing as expected. AI work deserves the same minimum standard once it moves beyond a single chat response.
Each card also streams the agent’s activity as it happens: the searches it runs, the tools it uses, the text it produces, and the files it writes. If one part of the run matters more than the others, you can scroll through that card and see the work that was done. You do not have to wait for the system to flatten the whole process into one polished output.
This is not an argument for micromanaging agents. Compound Console does not provide a general way to redirect an agent once it is running, and I would not want a system that assumes someone should supervise every step. The place to change the assignment is before execution, when the plan is still editable. During the run, the value of visibility is simpler: you can understand the state of the work without pretending that the final output is the whole story.
A Visible Worker Has a Role, an Assignment, and a Record
Seeing an agent work is useful in the moment, but it is only part of the picture. If an agent is going to be treated as a worker, the system also needs to preserve the context for judging its work afterward.
First, the agent catalog defines the specialist. In Compound, an agent definition is a durable description of what that specialist is and how it works: its role, standards, workflow, constraints, tools, and skills. It is not the exact one-off prompt for a particular task. A competitive-intelligence agent, for example, has a standing approach to competitive analysis that exists before this week’s request and will remain after it.
Second, the approved plan defines the assignment for this run. It shows which specialist was selected, what it was asked to produce, and what other work it depends on. That is the decision the human reviewed before execution began.
Third, each agent card retains its streamed activity and output in sequence for the run. After the run finishes, you can scroll back through that card to inspect how the output was produced. The final deliverable remains important, but it is an outcome. The card stream is the record behind it.
Those three layers establish different facts. The catalog tells you what kind of worker the agent was intended to be. The plan tells you what job it was given. The retained card stream tells you what it did. A final report cannot establish all three on its own.
Visibility Is Not the Same as Control
Live visibility has value even though it does not give you general mid-run control. It tells you whether work is active, what has finished, and where the process is waiting. It gives you a direct path into the work of a particular specialist rather than forcing you to infer the entire process from the final output. And it leaves a record for later review when the result deserves more scrutiny.
What it does not do is turn the human into a constant operator. That would recreate the same suboptimal relationship most people have with chat AI: the person has to sit beside the system, repeatedly supplying judgment and direction. The purpose of agent design and plan review is to place that judgment where it has the most leverage, before execution. Live visibility lets the person responsible maintain awareness while the system carries out the work.
The distinction matters because responsibility has two parts. Before the run, the human decides what work is authorized. During and after the run, the human needs enough visibility to understand what was done in their name. Neither requirement means reading every line. Both require more than accepting a polished output on faith.
The Record Behind the Result
Output can look excellent and still leave important questions unanswered. Did the research agent actually investigate the sources it was assigned? Did the strategist receive the inputs the plan said it would receive? Did a specialist stay within the role it was designed to play, or did the run drift into a different job because the task was poorly scoped? Those are important questions when the work informs a real business decision.
The catalog, plan, and retained card stream make those questions inspectable. You can compare the standing role against the assignment, then compare both against the record of the run. It gives the accountable person evidence to review instead of a conclusion to accept.
There is a tradeoff. More visibility creates more material to inspect, and nobody has time to audit every routine run in full. Most of the time, the status view and final deliverable will be enough. The retained card stream earns its place when the decision is significant, the result is surprising, or someone needs to understand why the system reached a particular conclusion. Good operations do not demand constant inspection. They make inspection possible when it is needed.
The final output is still the point of the work. But when agents are doing meaningful work, it should not be the only thing you can see. The human remains responsible for the outcome. A system that shows who is working, where the work stands, and how it was done gives that responsibility somewhere to go.