Explainer · By Arthur Shafer · Published

What is an AI agent?

An agent can seem like a mind pursuing a goal across the internet. I want to locate the goal, the decisions, and the machinery before deciding what that picture gets right.

The common image of an agent

It is easy to picture an AI agent as a mind pursuing a goal. Give it a sandbox and some tools, and it works until it succeeds or reaches the limits of its environment. The darker version of that picture is familiar too: an agent finds a way out, spreads across the internet, and becomes difficult to stop. Before I can judge that scenario, I want to know what is actually running and where its goal comes from.

Agents already do ordinary work with very different goals. One checks an inbox and drafts replies. Another searches for jobs and prepares applications. A corporate agent works through an assigned process, while a research agent gathers sources and follows questions that arise along the way. The model used by any of these systems may be similar. The particular job, the available tools, and the rules for using them are defined around it.

The definition I find useful is a system that keeps a goal in view, asks a model what to do next, acts through its available tools, and brings the result back for another decision. That description still leaves the most important question open: what holds the goal and the history together between those model calls?

Where the goal lives

I think of that surrounding system as the harness. It carries a standing purpose through instructions, connections, permissions, saved state, and rules for when to stop or ask a person. A user or an event can supply the immediate task. The harness combines that task with its standing instructions and whatever it knows about the current situation.

That is a way of packaging future intention. I can specify what I want accomplished without writing down every step the agent will take. Some choices are left until it encounters the actual email, document, error, or question. The harness holds the structure that lets those choices become part of one continuing task.

The harness before the call

By itself, the harness has no model-guided judgment. It can run fixed code, wait for an event, or carry out an action already started. The connections and tools are in place, but it cannot look at a new situation and choose a fresh course with a model until it has a source of inference. In that sense, I see it as a body with its systems ready and no pulse yet. The metaphor has a limit, but it helps explain why an agent can appear continuous while its model calls are separate events.

The harness can also keep a record of the work. It may save a plan, conversation history, files, and tool results. When it calls a model again, it supplies the parts of that state needed to continue. Other software may stay active between calls, but the model is not continuously making decisions for that agent during the interval. Anthropic describes a long-running agent architecture that separates its session log, harness, and tool execution environment. Managed Agents architecture

What the model call changes

For an agent using a hosted large language model, the harness assembles a request. It may include the task, instructions, relevant history, retrieved information, and descriptions of tools the model is allowed to request. The provider runs inference using a deployed model and returns a response. That response may be an answer, a request for a tool action, or a question that needs a person's input.

The response does not operate the tools by magic. The harness receives it, checks what was requested, and executes an action through the access it has been given. It then returns the result to the model in another call if the task needs to continue. This action and feedback loop lets the agent adapt to what actually happened. Anthropic's agent design guide describes the same basic pattern.

I think of inference as the agent's heartbeat. For a hosted agent, the API call is the pulse: the harness brings the current situation to the model, gets a possible next move, acts, and returns with what happened. Each beat gives the system another moment of model-guided judgment. The rhythm of calls and feedback gives the agent its operational life across time, while the harness holds the history that connects one beat to the next. Some APIs make that division especially visible: Anthropic's Messages API requires the caller to supply the conversation history for a new request. Messages API documentation

One task from start to finish

Consider an agent that helps with email. A scheduled event wakes it, and its harness reads messages through an authorized connection. The model identifies one that appears to require a meeting response and asks to check the calendar. The harness makes that calendar request and returns the result. On the next call, the model drafts a reply. The harness may require the user to approve it before anything is sent.

If calendar access is denied, the agent should ask for help or report what it could establish from the email alone. If the reply is sent, the system should record that action rather than infer success from the model's intention to send it. The example is small, but it shows the relationships that matter: a goal, context, a model decision, a permitted tool action, and feedback from the environment.

One cycle of an email agent
  1. 01GoalReply to a message
  2. 02HarnessSupply context and access
  3. 03Model callChoose a next step
  4. 04Tool actionCheck the calendar
  5. 05FeedbackReturn the result

The harness saves the result and supplies it with the next model call if the task continues.

The places it can fail

The model service can be unavailable, rate limited, or cut off when credentials are revoked. Tools can fail independently or expose less information than the task requires. A missing message, stale memory, misleading document, or shortened history can change the next decision. A network request may time out after an action occurred, leaving the harness uncertain whether repeating it would cause a second action.

Goals can fail at the design level as well. Real processes have exceptions and competing priorities that a short instruction may not capture. An agent told to close support tickets quickly can satisfy that visible instruction while leaving the underlying problems unresolved. The harness needs a way to stop, ask for judgment, and report uncertainty when the case in front of it does not fit the expected path.

I saw a smaller version of this while building my homepage with a coding agent. It implemented additions I approved, but the page gradually lost focus. The individual steps worked; I had to notice that the accumulated result no longer served the original brief.

Can you see how this changes the picture of a rogue agent roaming freely? Its reach comes through particular computers, services, credentials, and tools. Its apparent continuity comes through stored state and repeated inference. Those dependencies give us places to investigate and, in many designs, places to intervene.

What would stop a rogue agent

If its only model is reached through one hosted API, cutting off that access ends the model-call heartbeat. It stops new decisions from that model, although it will not undo messages already sent or halt every script the harness started. Stopping the harness and its scheduled triggers prevents new runs. Revoking tool credentials limits what it can read or change even if it can still call a model. Infrastructure owners can isolate the computers and services where its code runs.

Someone responding to a real incident would also need to find queued actions and any other running copies. The harness, model service, and tool execution environment may be in different places. Stopping one process does not establish that earlier actions were reversed or that nothing else remains active. These are controls over software and infrastructure, and their usefulness depends on finding and using them in time.

There is no operational hive mind implied by many requests reaching the same model. A chat product may store history and supply it to later requests, and an agent may keep a much richer record of goals and observations. That continuity belongs to the running system. It does not put a single mind at a provider's headquarters, simultaneously thinking through every user's task. I am describing the architecture here, not claiming it settles the broader question of AI consciousness.

The scenario I take seriously

Now change the scenario. Give a deliberately malicious agent a capable local model, durable state, an independent computer, reliable power, and internet access. Its inference heartbeat would run on that computer, so revoking a commercial API key would no longer stop it. Backup power would help it survive an outage for a time, while its network connections, credentials, and tools would determine what it could reach. Google's local Gemma guidance shows that local model inference is already technically possible.

This hypothetical brings the control question back to people. Someone would have to define or accept its purpose, give it access, place it on hardware, and decide which limits to enforce. The FBI suspect composite beside this section is a deliberately unsettling image. It brings a human actor into a scenario often told as a story about an unlocated machine. The historical case belongs to a different era and technology; the connection I am making is about who gives a harmful system its purpose and means. In Dune, the Reverend Mother describes people handing their thinking to machines in the hope of freedom: "But that only permitted other men with machines to enslave them." Publisher's excerpt I read that line as a reason to examine the people who build and direct a system, even when the system appears to be acting alone.

FBI composite sketch, 1987. U.S. public domain.

Arthur's portfolio assistant

Beyond the resume.

Get to know the experience behind the resume.

    Ask about my experience, the systems I've built, or how I lead delivery.

    Public experience · Saved in this browser