What an AI agent framework actually gives you
An AI agent is, at its core, a loop: call a model, let it decide whether to call a tool, execute the tool, feed the result back, and repeat until the task is done. You can write that loop in fifty lines against any model API. An AI agent framework is what you reach for when the fifty lines stop being enough — it supplies the reasoning loop, the tool-calling plumbing, state and memory handling, structured outputs, and some way to orchestrate multiple steps or multiple agents, so you are not rebuilding the same scaffolding for every project.
Concretely, most agentic AI frameworks bundle some combination of: a model abstraction so you can swap providers; a way to declare tools (usually from typed function signatures); a state object that persists across steps; a control-flow mechanism (a graph, a task list, a conversation, or plain function calls); and hooks for tracing and evaluation. Where they differ — and this is the whole point of this guide — is the programming model they impose on you.
One distinction is worth fixing before anything else: a framework defines how you write the agent; it is not the production runtime. Where the agent's process lives, how its tool calls are sandboxed, how it is scheduled, retried, and observed across thousands of runs — that is the job of an agent runtime, and every framework below still needs one underneath it. Choosing LangGraph or CrewAI answers "how do I express this workflow", not "how does it run safely at 3 a.m.".
The axes that actually matter when choosing
Feature matrices for ai agent frameworks age in weeks and mostly converge anyway — everyone supports tools, streaming, and the major model providers. The decisions that will still matter a year into the project are these:
- • Programming model. Do you want to express the agent as an explicit graph (LangGraph), a crew of roles with tasks (CrewAI), a conversation between agents (AutoGen/AG2), or type-safe functions (Pydantic AI)? This is the single biggest determinant of how the code reads six months from now.
- • Control vs abstraction. High-level frameworks get you a demo fast and then fight you when you need to intervene mid-run. Low-level ones make you wire more, but every transition is yours to inspect and change.
- • State and memory. Is state a first-class, typed object you can checkpoint, resume, and time-travel through, or an implicit message history? Long-running and human-in-the-loop workflows live or die on this.
- • Tool and MCP support. How tools are declared (typed signatures vs schemas) and whether the framework can mount MCP servers as tool sources, so you are not hand-wrapping every integration.
- • Observability. Can you trace every model call, tool call, and state transition out of the box, or do you bolt it on? You will debug agents far more than you write them.
- • Maturity and community. Age of the project, release cadence, and — as the AutoGen story below shows — whether the project's governance is stable.
- • Language. Python is the default across the board; LangGraph and LlamaIndex also ship TypeScript, and AutoGen has a .NET line. If your product is TypeScript-first, that narrows the field fast.
Keep those axes in mind as you read the framework profiles — each one is strong on some and deliberately weak on others, and the "best ai agent builder" for you is the one whose trade-offs line up with your workflow, not the one with the longest README.
The main open-source AI agent frameworks
LangGraph
LangGraph, from the LangChain team, models an agent as a graph: nodes are functions (model calls, tool executions, routing logic), edges define transitions — including conditional edges and cycles — and a shared, explicitly typed state object flows through the graph. Its persistence layer checkpoints that state at every step, which is what enables pausing for human approval, resuming after failure, and "time-travelling" to an earlier state to replay a run (LangGraph docs; LangGraph — Persistence). Because every transition is explicit, LangGraph is the strongest pick when you want control over a complex, stateful, possibly cyclic workflow — and the price is that you write that graph yourself rather than getting it for free from a higher-level abstraction. It ships for both Python and JavaScript/TypeScript (github.com/langchain-ai/langgraph).
CrewAI
CrewAI's unit of design is the crew: you define agents with a role, a goal, and a backstory, give them tools, assign tasks, and let the crew execute those tasks sequentially or under a manager. It is a role-playing, multi-agent-first programming model, and the appeal is speed — a researcher/writer/reviewer team is a few dozen lines. CrewAI also provides Flows for event-driven, more deterministic orchestration where you need explicit steps and state alongside crews (CrewAI docs — Agents; CrewAI docs — Flows; github.com/crewAIInc/crewAI). It is Python-only and is best when the problem naturally decomposes into roles and you want the team running quickly rather than hand-wiring every transition.
AutoGen and AG2
AutoGen originated in Microsoft Research and popularised the conversational multi-agent model: agents are "conversable", they exchange messages, group chats coordinate several of them, and a human can be a participant in the loop. Its programming model is still the cleanest way to think about agents that genuinely need to talk to each other (AutoGen docs).
The thing to know before you adopt it is that as of 2026 the AutoGen lineage has split and is consolidating. A group of the original maintainers created AG2, a community fork distributed as the ag2 package that keeps the classic autogen namespace and agent classes (ConversableAgent, GroupChat, and so on) available as "AG2 Classic" (AG2 docs; github.com/ag2ai/ag2). Meanwhile Microsoft's own AutoGen repository states that it is in maintenance mode and community-managed, and directs new users to Microsoft Agent Framework, which it describes as AutoGen's successor and for which it publishes a migration guide (github.com/microsoft/autogen; Microsoft — AutoGen to Agent Framework migration guide). The practical read: for a new project, treat AG2 as the community continuation of the conversational model and Microsoft Agent Framework as Microsoft's supported direction, and check the official repos for current status rather than relying on a 2024-era tutorial.
Pydantic AI
Pydantic AI comes from the team behind Pydantic, the validation library most Python AI code already depends on, and it is a type-safe, function-first framework: an agent is a typed object, tools are plain Python functions whose signatures become the tool schema, outputs are validated against Pydantic models (so a malformed response triggers a retry rather than a downstream crash), and a dependency-injection system passes connections, clients, or test doubles into tools and system prompts (Pydantic AI docs; Pydantic AI — Output; Pydantic AI — Dependencies). That combination makes it unusually pleasant to unit-test, and it is the natural choice for a Python team that already thinks in types and wants agents to feel like the rest of their codebase rather than a separate DSL. It is also younger than the others here, so check the docs and changelog for current API stability (github.com/pydantic/pydantic-ai).
| Framework | Programming model | Language | Best for |
|---|---|---|---|
| LangGraph | Graph of nodes and edges with explicit, typed state; supports cycles, branching, checkpointing | Python, TypeScript | Complex, stateful, branching or long-running workflows where you want explicit control |
| CrewAI | Role-based crews of agents assigned tasks; Flows for event-driven orchestration | Python | Standing up a role-playing multi-agent team quickly with minimal wiring |
| AutoGen / AG2 | Conversational multi-agent orchestration — agents exchange messages, group chats, human-in-the-loop | Python (AutoGen also .NET) | Conversation-style multi-agent research and prototyping; AutoGen itself is now in maintenance mode |
| Pydantic AI | Type-safe, function-first agents with Pydantic-validated structured outputs and dependency injection | Python | Python teams that value type safety, testability, and validated outputs |
| LlamaIndex (agents) | Agents and workflows built around a retrieval/indexing core | Python, TypeScript | Retrieval-heavy agents over your own documents |
Others worth knowing
- • LlamaIndex agents. LlamaIndex started as the indexing and retrieval library, and its agents and workflows are built around that core — the right starting point when the agent's main job is reasoning over your own documents (LlamaIndex docs).
- • OpenAI Agents SDK. Open-source and lightweight — agents, handoffs between agents, and guardrails — but it is a vendor SDK designed around OpenAI's platform rather than a framework-neutral layer (OpenAI Agents SDK docs).
- • Google ADK. Google's Agent Development Kit is likewise open source and model-agnostic in principle, but optimised for Gemini and the Google Cloud ecosystem (Google ADK docs).
- • Semantic Kernel and Microsoft Agent Framework. Semantic Kernel is Microsoft's long-standing SDK for integrating models into .NET, Python, and Java applications; Microsoft Agent Framework is the newer, consolidated direction Microsoft now points AutoGen users to (Semantic Kernel docs; Microsoft Agent Framework docs).
Get the weekly AI engineering brief
Agents, RAG, evals, and the tools worth using — one practical email a week. Plus the free roadmap PDF.
By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.
CrewAI vs LangGraph: the real trade-off
This is the comparison most teams are actually making, and it is cleaner than the marketing suggests. CrewAI optimises for abstraction and speed; LangGraph optimises for explicit control over state and flow. Neither is "more capable" — they sit at different ends of the same spectrum.
With CrewAI you describe who is on the team and what needs doing; the framework decides much of the how. That is wonderful on day one — a content pipeline, a research crew, or a multi-step analysis is up in an afternoon — and it is the right call when the problem really is "several specialists collaborating on a task" and you are happy to let the crew drive. The friction appears when you need to intervene mid-run, branch on a specific intermediate result, or guarantee a particular ordering: you end up reaching for Flows and writing the explicit steps anyway (CrewAI docs — Flows).
With LangGraph you write the graph. Every node, every conditional edge, every loop is yours, and the typed state that flows through it is checkpointed at each step, so pausing for a human decision, resuming after a crash, or replaying from an earlier checkpoint are supported behaviours rather than things you improvise (LangGraph — Persistence). The cost is that there is no "crew" to hand the problem to — the first version takes longer and the code is more verbose. That is a good trade when the workflow is genuinely complex, long-running, or has to be auditable; it is a poor trade for a weekend prototype.
A useful heuristic: if you can sketch your agent as a list of roles and tasks on a whiteboard, start with CrewAI. If you find yourself drawing boxes with arrows, loops, and "wait for approval here" annotations, you already want LangGraph. And the two are not mutually exclusive — teams do prototype in CrewAI to validate that the task decomposes well, then rebuild the production version in LangGraph once the shape of the workflow is known.
Decision guide: which framework for which situation
- • You want explicit control of a stateful, branching, or long-running workflow → LangGraph. Checkpointing and human-in-the-loop are built in, and the graph is auditable.
- • You want to stand up a role-based multi-agent team fast → CrewAI. Define roles and tasks, run the crew, iterate; graduate to Flows when you need explicit steps.
- • You are a Python team that values type safety and testing → Pydantic AI. Validated structured outputs and dependency injection make agents feel like the rest of your codebase.
- • Your agents genuinely need to converse with each other → AG2 for the community continuation of the AutoGen model, or Microsoft Agent Framework if you want Microsoft's supported path.
- • You are all-in on Microsoft, .NET, or an enterprise Azure stack → Semantic Kernel today, with Microsoft Agent Framework as the direction Microsoft is pointing to.
- • Your app is retrieval-heavy — agents reasoning over your own documents → LlamaIndex, so the agent layer sits directly on the indexing and retrieval you already need.
- • You are committed to one model vendor and want the thinnest layer → that vendor's SDK (OpenAI Agents SDK, Google ADK), accepting the lock-in.
- • One agent, a few tools, no multi-step state → no framework. A plain tool-calling loop against the model API is simpler to debug and has fewer moving parts.
Whichever you choose, remember that the framework is only half the problem. It gives you a way to write the agent; it does not give you a sandbox for its tool calls, a scheduler for its long-running jobs, durable storage that survives a redeploy, or the tracing and cost controls you need across thousands of runs. Those belong to the runtime layer — see agent runtimes explained for how that layer fits underneath any of the frameworks above — and the sooner a team separates "how we write agents" from "how we run agents", the fewer rewrites it faces later.
Frequently Asked Questions
What is an AI agent framework?
An AI agent framework is a library that gives you the scaffolding for building an agent: the reasoning loop that calls a model repeatedly, the plumbing for tool and function calling, state and memory handling, and a way to orchestrate multiple steps or multiple agents. Open-source examples include LangGraph, CrewAI, AutoGen (and its AG2 fork), and Pydantic AI. The framework defines how you write the agent; it is not the production runtime that sandboxes, schedules, and monitors it.
What is the best open-source AI agent framework?
There is no single best open-source AI agent framework — the right choice depends on the programming model you want and your use case. LangGraph is the strong pick when you need explicit control over a stateful, branching workflow; CrewAI is the fastest way to stand up a role-based multi-agent team; Pydantic AI is the choice for Python teams that prioritise type safety and testability; and LlamaIndex fits retrieval-heavy applications. Pick by how you want to express the agent's logic, not by feature-list length.
CrewAI vs LangGraph: which should I use?
Use CrewAI when you want a high-level, role-based abstraction — define agents with roles and goals, give them tasks, and let the crew run — and speed of getting started matters more than fine-grained control. Use LangGraph when your workflow has real branching, loops, or long-lived state and you want to define it explicitly as a graph of nodes and edges with built-in checkpointing. The trade-off is abstraction and speed (CrewAI) versus explicit control over state and flow (LangGraph).
Is Pydantic AI production-ready?
Pydantic AI is a newer agent framework from the team behind Pydantic, built around type-safe agents, Pydantic-validated structured outputs, and a dependency-injection system that makes agents straightforward to test. It is actively developed and used in production by teams who value those properties, but it is younger than LangGraph or AutoGen, so check the official docs and changelog for current API stability before committing a critical workload to it.
What happened to AutoGen?
AutoGen began as a Microsoft Research project for conversational multi-agent systems. As of 2026 the lineage has split: a group of original maintainers created the community fork AG2 (the ag2 package, which preserves the classic autogen namespace), while Microsoft placed the AutoGen repository in maintenance mode and now points new users to Microsoft Agent Framework, which it describes as AutoGen's successor and for which it publishes an AutoGen migration guide. Check the official AutoGen, AG2, and Microsoft Agent Framework docs for current status before starting a new project on any of them.
Do I need an agent framework at all?
Not always. For a single agent with a handful of tools, a plain tool-calling loop against the model's API — call the model, execute any tool calls it returns, append results, repeat — is often simpler, easier to debug, and has fewer dependencies. Frameworks earn their place when you need multi-step or multi-agent orchestration, durable state and checkpointing, structured outputs, or built-in observability, and when you would otherwise end up rebuilding those pieces yourself.
References
- • LangGraph — Official documentation
- • LangGraph — Persistence (checkpointing, human-in-the-loop, time travel)
- • LangGraph — GitHub repository
- • CrewAI — Official documentation
- • CrewAI — Agents concept
- • CrewAI — Flows concept
- • CrewAI — GitHub repository
- • AutoGen — Official documentation
- • AutoGen — GitHub repository (maintenance-mode notice)
- • Microsoft — Migrating from AutoGen to Microsoft Agent Framework
- • AG2 — Official documentation
- • AG2 — GitHub repository
- • Pydantic AI — Official documentation
- • Pydantic AI — Output (structured, validated outputs)
- • Pydantic AI — Dependencies (dependency injection)
- • Pydantic AI — GitHub repository
- • LlamaIndex — Official documentation
- • OpenAI Agents SDK — Official documentation
- • Google Agent Development Kit (ADK) — Official documentation
- • Microsoft — Semantic Kernel documentation
- • Microsoft — Agent Framework documentation