Pydantic AI is the Python AI SDK: a typed, extensible agent loop with every model a string swap away. The same agent runs everywhere you need it: behind a web frontend, in the terminal, on a voice call, on a durable background queue, or as a plain object you call run() on. Image generation and embeddings come in the same box.
Pydantic AI Harness has everything an agent needs for complex, long-running work, snapped on as capabilities, from memory, sub-agents, and context management to a complete coding agent.
View the complete documentation at pydantic.dev/docs/ai.
What are you building?
From simple typed data extraction to complex, long-running multi-agent collaboration, Pydantic AI and Pydantic AI Harness have got you covered.
Coding agent
A complete coding agent in your terminal: workspace-rooted file access, allowlisted shell, repo orientation, planning, and context management that survives long sessions. Here with web search and a second-opinion advisor snapped on alongside:
uv add pydantic-ai pydantic-ai-harnessfrom pydantic_ai import Agent
from pydantic_ai.capabilities import WebSearch
from pydantic_ai_harness import Advisor, Coder
agent = Agent(
'anthropic:claude-fable-5',
capabilities=[
Coder(), # files, shell, repo context, planning, sub-agents, context management
WebSearch(), # look up docs and error messages on the web
Advisor('openai:gpt-5.6-sol'), # a second opinion from another model when stuck
],
)
agent.to_cli_sync()Coder is a regular combined capability, not a black box: use it whole, or use the blocks it bundles directly; the two are equivalent:
capabilities = [
FileSystem('.'), Shell(cwd='.'), RepoContext(), Planning(), SubAgents(...),
ClearToolResults(), WarnNearLimits(), ToolOutputLimits(),
]Run the file and you're chatting with the agent in your terminal. To try it before writing any code, run the exported coder_agent with clai (the Pydantic AI CLI), via uvx:
uvx --with pydantic-ai-harness clai -a pydantic_ai_harness.coder:coder_agent -m anthropic:claude-fable-5Build this → Coder, from the Harness
Data extraction
Give the agent an output type and tools, and every run comes back validated and typed:
uv add pydantic-aifrom typing import Literal
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext
class Sentiment(BaseModel):
label: Literal['positive', 'negative', 'neutral']
score: float = Field(ge=-1, le=1)
agent = Agent('openai:gpt-5.6-sol', output_type=Sentiment)
@agent.tool
def recent_reviews(ctx: RunContext[None], product: str) -> list[str]:
"""Fetch recent review snippets for a product."""
return ['The new release fixed everything I complained about!']
result = agent.run_sync('How are people feeling about the Extract app?')
print(result.output)
#> label='positive' score=0.9The @agent.tool function receives a RunContext that carries your dependencies in; the rest of its signature and its docstring become the tool schema, arguments are validated before your code runs, and the run is guaranteed to return a Sentiment, so your IDE, type checker, and the LLM all agree on the returned type.
Build this → Agents, Function Tools, and Structured Output
Realtime voice
Put the same agent on a live voice session, tools and capabilities included:
uv add "pydantic-ai[openai-realtime]"import asyncio
from pydantic_ai import Agent
from pydantic_ai.capabilities import MCP
agent = Agent(
instructions='You are a helpful voice assistant.',
capabilities=[MCP('https://internal.example.com/mcp')], # capabilities work in voice too
)
@agent.tool_plain
def order_status(order_id: str) -> str:
"""Look up the status of an order."""
return f'Order {order_id}: shipped, arriving Thursday.'
async with agent.realtime('openai:gpt-realtime-2.1').session() as session:
microphone = asyncio.create_task(stream_microphone(session)) # chunks → session.send_audio()
speaker = asyncio.create_task(play_audio(session.stream_audio())) # model audio → your speaker
async for part in session.stream_transcripts():
print(f'{part.speaker}: {part.transcript}')The model calls your tools mid-conversation while it keeps talking, and every session is instrumented; voice is just another frontend, on OpenAI Realtime, Gemini Live, Azure, and xAI Grok Voice.
Build this → Realtime Voice
Durable background agent
Attach TemporalDurability and the same agent runs inside a Temporal workflow: every model and tool call becomes a durable activity, so a run working through a background queue survives restarts, failures, and long waits:
uv add "pydantic-ai[temporal]"from temporalio import workflow
from pydantic_ai import Agent
from pydantic_ai.capabilities import WebFetch, WebSearch
from pydantic_ai.durable_exec.temporal import PydanticAIWorkflow, TemporalDurability
agent = Agent(
'openai:gpt-5.6-sol',
instructions='Research the topic and write a structured brief.',
name='researcher',
capabilities=[WebSearch(), WebFetch(), TemporalDurability()],
)
@workflow.defn
class ResearchWorkflow(PydanticAIWorkflow):
__pydantic_ai_agents__ = [agent]
@workflow.run
async def run(self, topic: str) -> str:
result = await agent.run(f'Write a brief on: {topic}')
return result.outputDBOS and Prefect attach the same way, first-party and co-maintained, with Restate, Kitaru, and Airflow integrations besides.
Build this → Durable Execution
Image generation
Ask for an image and make it the run's typed output:
uv add pydantic-aifrom pathlib import Path
from pydantic_ai import Agent, BinaryImage
agent = Agent('openai:gpt-5.6-sol', output_type=BinaryImage)
result = agent.run_sync('Generate a minimalist logo for a coffee shop called Extract.')
Path('logo.png').write_bytes(result.output.data)Provider-native generation on models that support it (like this one), a subagent fallback you can configure for the rest, and a standalone image API on the way.
Build this → Image Generation
<!-- Embeddings section parked (bd54): restore by removing this comment. ### Embeddings Embed documents and queries for semantic search or a [RAG pipeline](https://pydantic.dev/docs/ai/examples/data-analytics/rag/): ```python from pydantic_ai import Embedder embedder = Embedder('openai:text-embedding-3-small') result = embedder.embed_query_sync('What is machine learning?') print(len(result.embeddings[0])) #> 1536 ``` Seven providers behind one typed API, [instrumented](https://pydantic.dev/docs/ai/integrations/logfire/) like everything else. It lives next to the agent that will use the results. **Build this →** [Embeddings](https://pydantic.dev/docs/ai/guides/embeddings/) -->Why Pydantic AI
-
Any model, one Python API. Virtually every model and provider (OpenAI, Anthropic, Google, Bedrock, Azure AI Foundry, Groq, Mistral, xAI, Ollama, and dozens more), swappable with a string, or through the Pydantic AI Gateway: one key for all of them, with failover and cost monitoring built in. No flagship feature is locked to one vendor.
-
…