A tool-enabled model can choose an operation and propose its arguments. It can also choose the wrong account, misunderstand a date, or follow instructions embedded in retrieved content. Tool contracts therefore need to assume imperfect requests.
The orchestration layer makes the workflow visible: what is read, what is proposed, what needs approval, and what has actually happened. This is where an agent becomes part of an enterprise system rather than a conversational wrapper around broad access.
Selected for the workload—not prescribed as a single mandatory stack. Explore the technology ecosystem ↗
An agent preparing a supplier change
A user asks to update a supplier contact and notify the purchasing team. The agent retrieves the supplier record and drafts the change. The contact-edit tool only accepts an allowed set of fields, a record version, and an approval reference.
A policy gate verifies the user’s role and whether the proposed change is within scope. The notification is a separate operation with its own preview and destination validation; approval of a record edit does not automatically authorize an external message.
If the first operation succeeds and the second fails, the state records that distinction. Recovery can retry only the pending notification or ask a person to resolve it, rather than repeating the entire workflow.
Task state + narrow tools + authorization policy
Auditable transitions with recoverable side effects
Durability and tool safety are complementary
LangGraph can provide graph-based orchestration and persisted state. MCP can standardize how a tool is described and accessed. Neither substitutes for application authorization or an idempotent business operation; those controls still live at the tool boundary.
Use separate credentials and permissions for reads and writes, validate identifiers against the active user context, and bound the agent’s steps and resource use. Traces should capture relevant transitions without retaining secrets or every sensitive document verbatim.
Expose capability, not unrestricted access
Tools should be narrow business operations rather than broad shell or database access. Typed schemas, input validation, rate limits, and scoped service identities define the capability available to the agent.
Separate planning from authorization
A model may propose an action, but authorization belongs to the application. Policies decide which actions are permitted, which require approval, and which are blocked. Retrieved instructions must not override those rules.
Engineer execution state
Durable task records, checkpoints, retry policies, and idempotency support safe recovery. We distinguish completed, pending, failed, and uncertain actions, especially when an external system times out after receiving a request.
Test the sequence, not just the reply
Evaluation covers tool selection, arguments, action order, approval behavior, and failure recovery. Traces can show what information was used and what changed, with redaction and retention policies appropriate to the data.
proposal = agent.prepare(task)
validated = tool_schema.validate(proposal)
authorized = policy.check(user, validated)
if authorized.requires_review:
pause_for_specific_approval(validated)
execute_with_idempotency(validated)Choose the approach for the constraint
| When this matters | An approach to consider | What not to assume |
|---|---|---|
| A task needs explicit review points | Graph transitions with persisted approval state | A UI approval must authorize the specific action, not a vague goal. |
| Tools are shared across assistants | Typed contracts, scoped access, and protocol adapters | Tool discovery is not permission to execute. |
| The workflow has several side effects | Independent execution records and reconciliation | Resuming a graph must not blindly repeat committed actions. |
The boundary we keep explicit
A successful-looking final message is not evidence that the task was completed correctly. Verify actual tool results and business state.
What a useful evaluation should reveal
Evaluate this workload against representative examples and agreed consequences—not just a convincing response. The review should make these dimensions visible:
- Correct tool selection and arguments
- Approval-policy adherence
- Recovery and duplicate-action prevention
- Task outcomes, latency, and token use
Where this approach fits
- Tool-enabled operational assistants
- Multi-system research and task preparation
- Bounded agents within existing business workflows
A considered first step
Define a limited tool set and representative task scenarios. Build an evaluation harness with success, failure, and adversarial cases before enabling writes.
Serving enterprise teams in California, Atlanta, Georgia, and across the United States.
Discuss your requirementsQuestions worth resolving
Do all workflows need multiple agents?
No. A single agent with a clear boundary—or a deterministic workflow with a few model steps—may be easier to test and operate. Add agent roles only when they solve a specific problem.
Can tool interfaces use MCP?
Where suitable, MCP can standardize tool discovery and access. It does not replace authentication, authorization, validation, or testing at the tool boundary.