Sufyan AbbasiA practical decision framework for choosing between single-agent and multi-agent LLM architectures — state management, failure handling, human oversight, latency, cost, reliability, and security, with version-aware comparisons of LangGraph, CrewAI, AutoGen, and Microsoft Agent Framework.
Teams adopting LLM agents face a recurring question: one agent with good tools, or several working together? Demos make the multi-agent answer look natural. Production is less flattering — coordination costs, state bugs, and evaluation burden arrive together, and late.
This article is a decision framework around the dimensions that determine whether multi-agent designs pay off: task structure, state management, failure handling, human oversight, latency, cost, reliability, and security. It compares LangGraph, CrewAI, and Microsoft AutoGen — plus Microsoft Agent Framework, now Microsoft's recommended choice for new projects.
Version note: Microsoft's AutoGen repository states AutoGen is in maintenance mode — no new features, community-managed — and recommends Microsoft Agent Framework (1.0, GA April 2026) for new projects. The comparison keeps AutoGen for existing systems and covers Agent Framework for new work. Claims reference October 2026 docs — verify before implementing.
Build the single-agent version first. One LLM loop with a defined toolset forces the hard questions early: task boundary, real tool needs, where the agent gets stuck.
Stay with one agent when:
None of these is automatic — each needs a concrete, checkable form. Treat them as hypotheses, not justifications.
Independent subtasks exist. Parallel research across unrelated sources or concurrent calls against different systems. The operative word is independent: if step B cannot start until step A finishes, you have a pipeline — and a pipeline does not necessarily need multiple agents either.
Roles genuinely differ in instructions, tools, or models. A researcher with web-search tools and a policy reviewer with compliance documents degrade each other when crammed into one system prompt. CrewAI models this directly: agents carry a role, goal, and backstory; tasks bind to agents with explicit expected outputs.
Isolation is engineered, not assumed. Splitting work across agents does not by itself contain failures — a shared downstream consumer still sees every upstream error. Fault isolation has to be designed: per-agent retry budgets, timeouts, validation at handoff boundaries.
The workflow is long-lived and interruptible. Processes spanning hours or days with human checkpoints need durable per-step state. LangGraph splits this into checkpointers (thread-scoped: continuity, human-in-the-loop, time travel, fault tolerance) and stores (cross-thread facts and preferences); Microsoft Agent Framework's workflow engine likewise offers checkpointing and hydration.
Trust boundaries differ across steps. When workflow parts operate under different credentials or data-access scopes, separate agents give each step a clean identity. Qualification: differing permissions do not always require multiple agents — one agent can assume scoped credentials per tool call. Multiple agents earn their keep when the boundaries are stable, auditable, and worth enforcing structurally.
Shared state is the actual complexity. Once agents must agree on shared facts — the current plan, the customer's intent — you have a distributed-systems problem. Conflicts and stale reads do not vanish because the components are language models; managing them is real engineering work.
Debugging surface multiplies. One agent, one trace. Five agents: reconstruct who said what to whom, in which order, on which version of shared state. Budget for observability or stay with one agent.
Latency stacks on the critical path. Sequential handoffs add model-call latency at every step. Count the sequential LLM calls before the user sees output; for interactive use cases this alone can disqualify a design.
Evaluation gets more complex. Testing interacting agents means covering the interaction space, not just individual outputs — demanding, not impossible. The workable pattern: per-agent unit tests plus integration tests on the composed workflow. If no one can say how the system would be evaluated, it is not ready to be built.
| Dimension | LangGraph | CrewAI | AutoGen (existing systems) | Microsoft Agent Framework (new projects) |
|---|---|---|---|
| Status | Actively developed | Actively developed | Maintenance mode — no new features; community-managed | 1.0 GA April 2026; Microsoft's recommended successor |
| Mental model | Explicit graph: nodes are steps, edges are transitions | Role-based crew: agents with roles, goals, backstories | AgentChat teams (0.4.x): predefined collaboration patterns | Agents + orchestrations + graph-based workflows; declarative YAML support |
| Coordination | You wire it: conditional edges, subgraphs | Declarative process: sequential (default) or hierarchical (requires manager_llm or manager_agent) | Round-robin, selector-based, swarm, graph flow teams | Sequential, concurrent, handoff, group chat, Magentic-One; workflow engine with branching and fan-out |
| State | Typed state schema; checkpointer (thread-scoped) + store (cross-thread) | Task outputs as context; structured TaskOutput; guardrails validate handoffs | Shared context within a team; memory components | Pluggable memory: conversational history, persistent key-value state, vector retrieval |
| Human-in-the-loop | interrupt(): pause anywhere, resume via Command | Per-task human_input flag | Approval steps within team design | HITL approvals, pause/resume for long-running workflows |
| Failure recovery | Resume from checkpoints; inspect historical state | guardrail_max_retries; manager re-plans in hierarchical mode | Termination conditions bound execution | Checkpointing; middleware hooks for safety/compliance |
| Best fit | Long-running, stateful, auditable workflows | Specialist teams with clear task boundaries | Maintaining existing AutoGen deployments | New enterprise projects, especially on the Microsoft stack |
Alt text: Flowchart for deciding between single-agent and multi-agent architectures. From "Task arrives": if the task is not decomposable into independent subtasks, use a single agent with well-chosen tools. If decomposable but roles, models, or trust boundaries do not differ, use a single agent with parallel tool calls. If they differ and the workflow is long-lived, interruptible, or auditable, use graph-based LangGraph. Otherwise, if starting new work on the Microsoft stack, use Microsoft Agent Framework; if manager-led delegation is needed, use hierarchical CrewAI; otherwise this path applies to existing AutoGen systems only (maintenance mode).
Alt text: Three multi-agent coordination patterns. Sequential: Task 1 (Agent A) flows to Task 2 (Agent B) flows to Task 3 (Agent C). Hierarchical: a manager agent that plans and validates delegates to Worker 1 and Worker 2, which return results to the manager. Team chat: Agent 1, Agent 2, and Agent 3 all read from and write to a shared team context.
Do not mix API generations: AutoGen 0.2's GroupChat/GroupChatManager/max_round/TERMINATE-token idiom is superseded, and even the 0.4.x AgentChat API is now in maintenance mode. For new projects Microsoft points to Agent Framework, which unifies Semantic Kernel's enterprise foundations with AutoGen's orchestration research.
Three patterns:
Explicit state (LangGraph). You declare the schema; nodes read and write named fields. Two persistence systems with different jobs: the checkpointer saves thread-scoped snapshots (pause/resume, time-travel inspection); the store keeps cross-thread facts. Two facts teams get wrong: MemorySaver/InMemorySaver keep checkpoints in RAM and lose everything on process restart — production needs a durable checkpointer — and resuming from an interrupt() restarts the entire node from its beginning, re-executing earlier code. Side effects before an interrupt must be idempotent, or moved after it.
Validated handoffs (CrewAI). Task outputs flow downstream as context, wrapped in structured TaskOutput (raw, JSON, or Pydantic-validated). The corrective to the "no inspectable state" caricature: guardrails. A task can carry function-based or LLM-based guardrails that validate outputs before the next task runs, with configurable retries (guardrail_max_retries) — a concrete mechanism for enforcing state invariants at handoff boundaries.
Conversation as state (team-chat architectures). Agents coordinate through shared conversational context — expressive, but hardest to reason about. State invariants are still enforceable, only more expensive: validation steps, structured outputs, termination conditions that halt the team on violation.
Write the invariants down first, whichever pattern you choose: which facts must never be stale, who may change them, what halts the system on disagreement.
Assume every component fails; the design question is blast radius:
Enterprise use makes this non-optional, and the frameworks differ in granularity:
interrupt() pauses anywhere — inside a node, even inside a tool — persists state via the checkpointer, and resumes via Command. Static interrupt_before/interrupt_after cover known checkpoints. Remember the re-execution rule: keep pre-interrupt code side-effect-free or idempotent.human_input flag puts human review at task boundaries — coarser-grained, simpler to reason about.Gate where the cost of being wrong exceeds the cost of waiting: production writes, customer-facing sends, financial actions. Drafting an internal summary does not earn a gate.
Latency: count sequential LLM calls on the critical path; multiply by typical per-call latency. Redesign before building if the total exceeds the UX budget.
Cost: (calls per run) × (tokens per call) × (runs per day) × (price per token). Multi-agent designs multiply the first three factors; run the arithmetic with real volumes, not demo volumes.
Reliability: the textbook estimate — ten sequential steps at 95% each yield about 60% end-to-end (0.95^10 ≈ 59.9%) — rests on a simplifying independence assumption (uncorrelated failures). Correlation can shift that estimate in either direction depending on the dependence structure: shared failure causes can make joint failure more likely, while shared mitigations (common verification, retries) can make it less likely. Use the number as a starting argument for adding verification and measuring your own step reliabilities, not as a prediction.
Security: more agents mean more trust boundaries. Scope credentials per agent or per tool call; treat inter-agent messages as untrusted input across trust levels (prompt injection propagates through agent teams as through microservices); log decisions for audit. One over-privileged agent is a single point of risk; five credential sets are five configurations to get wrong.
A short pilot against the single-agent baseline, measured on latency, cost, and reliability, teaches more than any diagram.
Hypothetical illustration — not a real deployment.
A support-triage workflow: tickets are categorized, enriched with customer history, and routed, with a human approving SLA-changing routing.
"It feels more scalable" is not a signal. A measured baseline is.
Disclosure: This article was created with AI assistance. Technical claims were checked against official framework documentation in October 2026; verify against current documentation before implementing.
Proposed author: Sufyan Abbasi — IT Expert, Bitneka Technologies (13 years overall IT experience; AI and LLM technologies). Draft prepared for the author's review and approval; it should not be presented as the author's writing unless reviewed, verified, and approved.