The agent needs more than a prompt.
Tyler Heshka · https://heshka.com/work/casebook/agent-workflows
- System Trace
- Applied AI · Agent workflows
- 4 min read
Giving AI agents substantial work, with enough context to proceed and clear limits on what they can decide.
My part
Task direction, context and instruction design, evidence review and approval.
01
Give it enough context to work.
I use AI regularly for software engineering, research, documentation and working through source material. Often I start in a chat tool, using it to plan the work and prepare instructions for an agent. The agent may help research, draft, implement and test. I review what comes back and decide what needs changing.
For a substantial task, I make the objective, relevant sources, settled decisions and limits explicit. I also describe what a useful result needs to prove. If the work is too broad, I break it into smaller tasks or separate the research from the implementation.
I don’t want a long conversation to be the only place the project remembers its decisions. I keep evidence, plans, approvals and current state in durable records. Reusable ways of working go into playbooks, skills or operating instructions. A fresh session can read those records and check the current source before continuing.
02
Decide what it is allowed to decide.
An agent can carry out substantial work without being given authority over every decision around it. I can authorize an implementation while keeping product and architecture decisions open for review. Permission to change code does not automatically include permission to commit, push or deploy it.
The same applies to information. I give the tool the context the task needs and leave out sensitive details that do not help it solve the problem. Where the material is confidential, I may generalize it, withhold it or use a different workflow. Removing a name does not necessarily make the rest safe to share.
I remain responsible for the factual claims and publication decisions I approve. Decisions that belong to a client, employer or other decision-maker still belong to them.
03
Make the result prove itself.
A clear explanation helps me understand a proposal. I still need to check whether it agrees with the evidence. That can mean going back to a source document, inspecting the current code or researching a claim independently.
For software work, I require tests, validation and review before committing. A passing test supports the behaviour it actually checks; it does not prove every claim in the agent’s summary. I also compare the result with the agreed scope and the decisions it was meant to preserve.
The checks depend on the work. A factual statement needs a source. A code change needs relevant tests. Public wording needs review, and publication or a release needs its own permission. Fluent output does not remove any of those responsibilities.
04
Correct the process, not just the answer.
While developing this site, I had to correct descriptions that made a development experiment sound like a production implementation. I also had to separate the order in which something happened from the procedure we intended to recommend. In both cases, a plausible account blurred a distinction the evidence needed to preserve.
When that kind of mistake repeats, I look at the context and instructions around the task. Was the source clear? Had I separated a hypothesis from a settled decision? Did the approval boundary need to be more explicit? I update the relevant records, instructions or checks so the correction is available to the next session too.
I take a similar approach when an agent gets stuck: clarify the context and task first, then consider a different model or tool if the work needs a capability the current one lacks. I expect mistakes. What matters is being able to catch them, understand them and change how the next piece of work is handled.
Work, review and revision
- Goal
- Context & boundaries
Before sharing: consider sensitivity and limit the context
- Agent work
- Result & evidence
- Review & checks
Factual claims → source checks
Code changes → relevant tests
Publication or release → specific approval
- Authorized change
Article link: https://heshka.com/work/casebook/agent-workflows
