Method

The Anatomy of an Agent

Answer capsule

Industry cannot agree on what an AI agent is, and Gartner predicted in June 2025 that over 40% of agentic AI projects will be cancelled by the end of 2027. This piece argues those failures trace to missing anatomy, and sets out the nine planes Red First declares for every agent: Soul, Knowledge, Skills, Capabilities, Listeners, Connections, Surfaces, Heartbeats, Guardrails. Whatever a build is called, it answers all nine questions somewhere, the only choice is whether they are answered on purpose.

What IS an agent?

A business owner shopping for an AI agent in 2026 is buying a word with no settled meaning. TechCrunch put it on record in March 2025: the biggest companies in the industry define "agent" in ways that contradict each other, and sometimes contradict themselves S1. Ryan Salva, a senior director of product at Google, told the publication the industry "overuses the term 'agent' to the point where it is almost nonsensical" S1.

The contradictions run inside single companies. In one week of March 2025, OpenAI published a blog post defining agents as automated systems that independently accomplish tasks, and developer documentation defining them as language models equipped with instructions and tools S1. Microsoft draws a line between agents and assistants. Salesforce lists six separate categories of agent on its own website S1.

Anthropic drew the sharpest line available in December 2024. Workflows follow code paths a developer fixed in advance; agents "dynamically direct their own processes" and choose their own route to the goal S2. The same essay advises builders to start with the simplest system that works, which often means no agent at all S2. It is a useful distinction, and it still describes behaviour, which means it describes a matter of degree.

Researchers have largely declined to settle the question. A Princeton team noted in July 2024 that under the classical textbook definition a thermostat qualifies as an agent, treated agency as a spectrum, and refused to add yet another definition to the pile S5. Others have replaced the binary with ladders: one 2025 framework defines five levels of agent autonomy by the role the human keeps, from operator down to observer S6. Vicent Botti argued in June 2025 that "agentic AI" is a rebadging of what thirty years of agent research already defined, with the prior work largely ignored S7.

The confusion has a price, and the price has a date. On 25 June 2025 Gartner predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, naming three causes:

  • escalating costs
  • unclear business value
  • inadequate risk controls S3.

The same release described "agent washing", the rebranding of assistants, chatbots and process automation as agentic AI, and estimated that only about 130 of the thousands of vendors claiming agentic capability are genuine S3. Anushree Verma, the Gartner analyst behind the forecast, said most such projects "are early stage experiments or proof of concepts" driven by hype S3.

A year on, Forbes revisited the forecast and found it aging well, adoption is broad but production use is thin, and much of what was counted as agentic was never agentic to begin with S11. Jim Rowan, head of AI at Deloitte, had already named the mechanism in the TechCrunch piece. Without a standardised definition it becomes hard to benchmark performance or keep outcomes consistent S1. Buyers are evaluating against a word. The word does not hold still.

A job description for software

The confusion when examined has a consistency to it. Every definition above describes behaviour. The thing acts on its own, pursues goals, uses tools. Behaviour comes in degrees, so behavioural definitions blur, and a chatbot with a loop can wear the label. An anatomy describes parts. A part is either present or absent, and that difference can be checked before a contract is signed.

The field had the right starting point thirty years ago. Stuart Russell and Peter Norvig defined an agent in 1995 as anything that perceives its environment through sensors and acts on it through effectors S4. Michael Wooldridge and Nick Jennings refined it the same year with four properties: autonomy, social ability, reactivity, and pro-activeness S4. Those four are still cited as the cornerstones of artificial agency S7. What never followed was the parts list: the set of organs a working system needs so those properties show up on purpose rather than by accident.

Recent research is converging on partial lists. One line of work describes modern agents as compound systems: a foundation model plus the scaffolding that gives it planning, memory and tool use S9. A January 2026 survey decomposes agents into six architectural dimensions S8. A 2026 developer guide runs seven cognitive functions, and is one of the few schemes anywhere to include governance as a function at all S10 S8. The lists overlap, none of them match, and the model itself is only one entry on each. The industry knows a parts list is needed. It has not finished one.

Red First's definition of an agent is occupational, a member of staff made of software. You would never hire a person on the strength of "autonomous and goal-directed". You would ask who they are, what they know, what they have been trained to do, what they may touch, and who they answer to. We build every agent to a fixed anatomy of nine planes, and each plane is one of those questions asked precisely. A blank answer is still an answer - it means the decision was made by accident.

The occupational definition can be tested against something every business already writes. Set the nine planes beside the sections of a well-written job description and the columns line up.

  • Soul - Job title, purpose of the role, standards of conduct
  • Knowledge - Required knowledge and qualifications, and what the holder is expected to learn in post
  • Skills - The skills and competencies section
  • Capabilities - Systems access and authority granted: what the holder may do and approve
  • Listeners - What the role monitors: the inbox, the queue, the accounts under watch
  • Connections - Key working relationships, and the accounts and credentials issued at onboarding
  • Surfaces - Where the work happens: the meetings attended, the reports delivered, the channels used
  • Heartbeats - The recurring duties: the Friday report, the month-end reconciliation
  • Guardrails - Limits of authority, escalation rules, compliance obligations

A business that can write a job description already knows how to specify an agent. The anatomy asks nothing new; it asks the same questions of a different kind of staff.

Nine plane cards arranged in an orbit around a single agent in the Red Command interface
All nine planes of one agent, declared on one screen. Demo tenant only.

1. Soul. Who is this agent, and what does it stand for? Identity, working doctrine, tone. This is the job description, written before any tool is granted. A credit-control agent and a customer-success agent can hold the same accounting access and be opposite members of staff: the first is built to collect firmly and protect the ledger, the second to spot trouble early and protect the relationship. In marketing, the Soul is where the brand's voice and its banned phrases live, so every draft the agent produces starts inside the house style instead of being corrected into it.

2. Knowledge. What does it know, and does it remember? The agent's knowledge base and its memory. Memory is the most loosely sold part of the market: session history is routinely marketed as memory, and a worker who forgets everything each night is not accumulating value. A service-desk agent that remembers a customer's last three issues opens the fourth conversation ahead instead of from zero. An HR onboarding agent that knows the handbook and remembers where each new starter has got to chases the one missing document rather than re-sending the full pack.

3. Skills. What has it been trained to do? Named, repeatable procedures, as distinct from raw tool access. A finance agent's month-end pack is a Skill: a written procedure covering what to gather, in what order, and what the output must contain. A sales agent's quote follow-up sequence is another. Two agents can hold identical tool access while only one knows either procedure, which is why "give it more tools" is the standard fix for an agent that was never trained for the job. A procedure is training. A tool is equipment.

4. Capabilities. What may it act with? The declared list of tools and services the agent can invoke: its hands. An invoicing agent might be able to raise draft invoices in the accounts system and nothing else: no credit notes, no payment runs. A procurement agent might raise supplier orders below a set value, with anything above routed to a person. This is also the first agent-washing test. A vendor selling a genuine agent can produce the list; a chatbot has no such list, because it touches nothing.

5. Listeners. What does it perceive? The signal sources the agent watches: the sensor half of the 1995 definition S4, and the half most products shipped without. A sales agent watching the shared enquiries inbox notices the lead that arrived at six on Friday evening. An operations agent watching stock levels sees the reorder point crossed before the fitter finds the shelf empty. A system that perceives only what you type at it is reacting to you, not to the business.

6. Connections. What does it reach into, under whose credentials? The systems the agent connects to, and the identity it carries when it does. A bookkeeping agent reaches into the accounting package under its own named identity, so its entries are distinguishable from a person's and its access can be revoked in one place. A scheduling agent connects to the company calendar with rights to read and propose, and nothing wider. Credential sprawl is where risk controls die in practice, and this plane is where they live.

7. Surfaces. Where do people work with it? The places a human and the agent meet. A finance director might meet the month-end agent in a review panel showing drafts against approvals. A warehouse team might meet the stock agent in a morning pick list on a wall screen. Technical taxonomies mostly describe agents as headless processes, yet installations fail on people far more often than on configuration. The same agent with no surface is still running, but nobody is working with it.

8. Heartbeats. When does it act unprompted, exactly? The declared schedules on which the agent moves without being asked. A cash-position summary at seven each working morning is a heartbeat. So is a Friday sweep of the pipeline for stalled deals, and a first-of-the-month scan for contracts entering their notice window. This plane dissolves the autonomy fog: outside, autonomy is a marketing adjective on a spectrum nobody can measure; here it is a timetable you can read, and every unprompted action traces back to a line on it.

9. Guardrails. What bounds the other eight? Scope, disclosure rules, and who may change the agent itself. A credit-control agent may chase, remind and reconcile, and must hand any settlement negotiation to a named person. A marketing agent may draft the week's posts and may not publish one; publication needs a human approval every time. Governance appears in almost none of the published taxonomies S8 S10, and Gartner's "inadequate risk controls" is this plane missing in the wild S3. An unbounded agent is a private experiment. A bounded one can be a colleague.

Map the nine back to 1995 and the fit is exact. Reactivity is Listeners. Pro-activeness is Heartbeats. Social ability is Connections and Surfaces. Autonomy is Heartbeats operating inside Guardrails. Wooldridge and Jennings named the properties a real agent exhibits S4; the planes are the organs that produce those properties on purpose. Nothing here overturns the academic work. It finishes the job the definitions started.

Unasked questions get answered in production

The nine questions work as a test before they work as a design. Take Gartner's three causes of cancellation and ask which planes were blank S3.

  • Unclear business value means Soul and Knowledge went undeclared: no job description, no memory, so nothing compounds and nothing can be measured.
  • Inadequate risk controls means Guardrails and Connections went undeclared: no bounds, credentials wherever someone pasted them.
  • Escalating costs means Capabilities and Heartbeats went undeclared: an unbounded set of actions running on an unbounded schedule.

Read this way, the projects did not fail in production. They failed at the description stage, they just didn't know until the time was wasted and the budget reported it later.

The same questions expose agent washing in a single sales call. Ask the vendor three things. Which systems does the product watch without being prompted? What does it do on a schedule, and where is that schedule written down? What bounds it, and who can change those bounds? A rebranded chatbot has no answer to any of the three, whatever the deck says. Gartner's estimate that roughly 130 of thousands of vendors are genuine S3 suggests how often those questions would end the meeting.

The deeper point is that the questions never go away. Every deployment answers all nine, whether or not anyone asks them. Skip Guardrails, and your bounds are whatever the model happens to refuse on a given day. Skip Heartbeats, and your autonomy is whatever the loop happens to do. Skip Soul, and your agent's identity is the last edit somebody made to a prompt file. Unasked questions get answered by accident, in production, at full price.

None of this requires our labels. Call the planes anything you like, or nothing at all. If you build something and call it an agent, the nine decisions are yours either way; the only choice available is whether you make them on purpose, in writing, before the thing goes to work. Writing them down first is the entire discipline. We named ours because a named decision can be reviewed, audited and improved, and an unnamed one cannot.

Get the nine answers in writing

For a business owner, the anatomy is a buying tool before it is anything else. Before signing for anything sold as an agent, ask for the nine answers in writing: identity, knowledge and memory, trained procedures, permitted actions, watched signals, connected systems and credentials, working surfaces, schedules, and bounds. A genuine vendor can produce them in a page. The 40% cancellation forecast S3 is, in large part, a queue of buyers who never got that page.

For a build, the anatomy turns a project from a hope into a specification. Value becomes measurable because Soul defines the job and Knowledge lets results accumulate against it. Risk becomes governable because Guardrails and Connections are declared rather than discovered. Cost becomes predictable because Capabilities and Heartbeats set the envelope in advance. The three Gartner failure causes are the same three properties, missing.

The pattern behind most failed installations is ungoverned, single-player AI - one person, one prompt box, no shared memory, no declared bounds, no record of what was done in the company's name. The anatomy is the opposite of that pattern by construction, because eight of the nine planes describe things a private prompt box does not have. Every agent we ship inside RedOS carries all nine planes declared before it starts work, and the declaration is visible on one screen.

As of June 2025, Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, and estimates that only about 130 of the thousands of vendors claiming agentic AI are genuine S3. The gap those numbers describe is anatomical: the market sells behaviour while nobody checks for parts. An agent is nine questions with written answers, and a build that cannot produce the answers has already told you what it is. What one full set of answers looks like is below.

One credit controller, written in full

Here is the anatomy applied to one agent a small business might run - credit control for a firm invoicing on thirty-day terms. It is an illustration of the declaration. It's not real, simply an example, no client sits behind it, and it makes no performance claims. The point is that the whole specification fits on a page.

Soul. The company's credit controller, in software. Its purpose is to keep cash arriving on time without damaging trading relationships. Its register is courteous, specific and firm. It never apologises for asking to be paid on the agreed terms, and it never threatens.

Knowledge. The company's payment terms and escalation policy, each customer's invoicing history and payment pattern, and a memory of every past exchange about every invoice, so the third reminder knows what the first two said and to whom.

Skills. The reminder sequence: a note seven days before an invoice falls due, on the due date, and at seven, fourteen and twenty-one days overdue, each written to its own register. The statement-preparation procedure. The weekly aged-debt summary procedure.

Capabilities. It may read the sales ledger, send reminders from the credit-control mailbox, prepare customer statements, and flag accounts for attention. It may not issue credit notes, alter an invoice, agree a payment plan, or place an account on stop.

Listeners. It watches the sales ledger for invoices crossing each ageing threshold, and the credit-control mailbox for replies.

Connections. The accounting package, under its own named identity, with rights to read the ledger and write nothing beyond notes. The email system, for the one mailbox it sends from. Both revocable in one place.

Surfaces. A weekly aged-debt panel the owner reviews on Monday mornings, and a thread per chased account showing every message sent and every reply received.

Heartbeats. A ledger scan at eight each working morning. Reminders dispatched on the schedule each invoice sets. The aged-debt summary every Friday. The statement run on the first of the month.

Guardrails. It escalates to the owner the moment a customer disputes an invoice, asks for a payment plan, passes thirty days overdue, or replies with anything it cannot classify. It identifies itself as an automated assistant of the company when asked. Only the owner may change any of the nine entries above.

Nine entries, one page, written before the agent sends its first reminder. This is the page the essay's close describes: written answers, produced in advance, that make the word "agent" mean something you can check.

Sources

  1. S1 Tier 2 · secondary press-article
    No one knows what the hell an AI agent is
    TechCrunch · Maxwell Zeff, Kyle Wiggers · 14 March 2025
    overuses the term 'agent' to the point where it is almost nonsensical
    Supports Industry-wide definitional contradiction: Salva quote, OpenAI blog vs docs, Microsoft agent/assistant split, Salesforce six categories, Deloitte on benchmarking
  2. S2 Tier 1 · primary company-disclosure
    Building effective agents
    Anthropic · Erik S., Barry Zhang · 19 December 2024
    dynamically direct their own processes
    Supports Workflow vs agent architectural distinction; advice to build the simplest system that works
  3. S3 Tier 1 · primary analyst-press-release
    Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027
    Gartner · Gartner Newsroom (analyst: Anushree Verma) · 25 June 2025
    are early stage experiments or proof of concepts
    Supports 40% cancellation forecast by end-2027; three causes; agent washing; ~130 genuine vendors of thousands
  4. S4 Tier 1 · primary research-paper
    Fully Autonomous AI Agents Should Not Be Developed
    arXiv (2502.02649) · Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, Giada Pistilli · 2025-02
    Supports Compiled canonical definitions: Russell & Norvig 1995 sensors/effectors; Wooldridge & Jennings 1995 four properties (autonomy, social ability, reactivity, pro-activeness)
  5. S5 Tier 1 · primary research-paper
    AI Agents That Matter
    arXiv (2407.01502) · Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, Arvind Narayanan · 2024-07
    Supports Thermostat qualifies under the traditional definition; agency treated as a spectrum; authors decline to add a new definition
  6. S6 Tier 1 · primary research-paper
    Levels of Autonomy for AI Agents
    arXiv (2506.12469) / Knight First Amendment Institute · K. J. Kevin Feng, David W. McDonald, Amy X. Zhang · 2025-06
    Supports Five levels of agent autonomy defined by the role the human keeps, operator through observer
  7. S7 Tier 1 · primary research-paper
    Agentic AI and Multiagentic: Are We Reinventing the Wheel?
    arXiv (2506.01463) · Vicent Botti · 2025-06
    Supports Agentic AI as rebadging of established agent research; the four 1995 properties still cited as cornerstones of artificial agency
  8. S8 Tier 1 · primary research-paper
    Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
    arXiv (2601.12560) · Arunkumar V, Gangadharan G.R., Rajkumar Buyya · 2026-01
    Supports Published taxonomies decompose agents into differing capability dimension sets, with governance largely absent
  9. S9 Tier 1 · primary research-paper
    The AI Agent Index
    arXiv (2502.01635) · Stephen Casper, Luke Bailey, Rosco Hunter, Carson Ezell, Emma Cabalé, Michael Gerovitch, Stewart Slocum, Kevin Wei, Nikola Jurkovic, Ariba Khan, Phillip J. K. Christoffersen, A. Pinar Ozisik, Rakshit Trivedi, Dylan Hadfield-Menell, Noam Kolt · 2025-02
    Supports Modern agents as compound systems: foundation model plus scaffolding for planning, memory and tool use
  10. S10 Tier 3 · commentary project-blog
    Types of AI Agent Architectures: 2026 Developer Guide
    MLflow · MLflow project · 1 July 2026
    Supports Published taxonomies decompose agents into differing capability dimension sets, with governance largely absent
  11. S11 Tier 2 · secondary press-article
    Why 40% Of Agentic AI Projects May Be Canceled By 2027
    Forbes · Robert J. Szczerba · 7 July 2026
    Supports One-year retrospective on the Gartner forecast: broad adoption, thin production use, much counted work never agentic

If you want the nine questions asked of your own business, a Red Brief is where that starts.