adk-python: Google’s Code-First Agent Runtime

What the Runner, the workflow scheduler, and the session services actually do once you read past the README.

Repository: google/adk-python · Primary language: Python · Size: 2,776 files / 744,143 LOC · License: Apache-2.0 · Stars: 21,701 · Analysed at: 9f7965e (2026-10-02)

Every team that ships an LLM agent eventually rewrites the same plumbing: conversation state, tool dispatch, confirmation gates, resumability after a human pauses the run, and some way to check the thing still works. adk-python is Google’s attempt to make that plumbing a library instead of a recurring project — a Python framework for defining agents and graph workflows in code, running them against Gemini or other models, and deploying them to Cloud Run or Vertex AI Agent Engine. It is Apache-2.0, Python 3.10+, roughly 21.7k stars and 41 contributors, with bi-weekly releases and a large dependency surface. By the end of this piece you should understand what the Runner actually does per turn, how the workflow scheduler decides between replay, rehydration, and fresh execution, and where the session, artifact, and memory services sit relative to the agent code you write.

What adk-python is, and what it is not

ADK is a Python library and CLI for defining LLM agents and graph workflows in code, running them against Gemini or other models with tool calling, session state, and human-in-the-loop pauses, and deploying them to Cloud Run or Vertex AI Agent Engine. It provides the runtime and the service interfaces; you provide the instructions, tools, and topology.

It is not a hosted agent platform, not a prompt-management or observability product, and not a model. It is also not a thin wrapper over the Gemini SDK. The Runner, the workflow graph, the session services, and the plugin system are the substance; the model call is one step inside them.

In a stack diagram it sits above the model SDKs (google-genai, LiteLLM, Anthropic) and below your application or your deployment target. It owns orchestration, state, and tool dispatch. It delegates inference to a provider and persistence to whichever session, artifact, or memory service you configure.

Maturity signals are mixed. The repository shows 21.7k stars, roughly 41 contributors, and 150 commits in the sampled window, with bi-weekly releases at v2.11.0. At the same time, many features are explicitly marked experimental or deprecated in their docstrings, so the API surface is still moving.

The audience is Python developers and platform engineers who want a supported runtime for multi-agent systems, tool calling, HITL, and deployment, and who accept a large dependency surface and a fast-moving API in exchange.

Architecture: four layers and one run loop

ADK decomposes into four tiers. Entrypoints — the Click CLI (adk run, adk web, adk eval, adk deploy), the FastAPI dev server, and the A2A executor — translate argv, HTTP requests, or A2A messages into Runner calls. The control plane holds Runner, App, PluginManager, evaluation, and telemetry configuration; it decides which node runs and what gets persisted, but never calls a model or a tool itself. The runtime holds BaseAgent/LlmAgent, Workflow and DynamicNodeScheduler, the LLM flows, and tools; it executes the graph and produces Events, reaching outward only through BaseLlm, BaseTool, and the service interfaces. Transport and storage — GeminiLlmConnection, the MCP session manager, GCP clients, and the session/artifact/memory/credential services — is the only tier that touches the network or disk.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
flowchart TD
    CLI["Entrypoints<br/>CLI · FastAPI dev server · A2A executor"]
    CP["Control plane<br/>Runner · App · PluginManager · eval · telemetry"]
    RT["Runtime<br/>LlmAgent · Workflow/DynamicNodeScheduler · LLM flows · tools"]
    TS["Transport & storage<br/>GeminiLlmConnection · MCP sessions · GCP clients · session/artifact/memory/credential services"]

    CLI -->|"Runner.run_async"| CP
    CP -->|"BaseNode.run_async"| RT
    RT -->|"BaseLlm.generate_content_async"| TS
    CP -->|"get_session / append_event"| TS

Runner is the single orchestration entry point. Its __init__ docstring states that exactly one of app, agent, or node must be provided, that agent or node gets wrapped into an App internally, and that providing app is the recommended way. The older agent=/node= paths still exist and must stay behaviorally identical.

One turn proceeds as follows. Runner.run_async loads the session through the session service and builds an InvocationContext. find_agent_to_run picks the node: a Workflow root routes itself, a resumable session routes to the author of the last function call, and otherwise the last transferable agent in the event history wins, falling back to root. The LLM flow assembles an LlmRequest from session events, the agent instruction, and canonical tools. The model returns text or function calls; calls are matched to tools and executed. Each Event is stamped with branch and isolation scope, merged with run-level custom metadata, yielded to the caller, and appended to the session.

DynamicNodeScheduler is the single driver for both workflow-integrated dynamic execution and standalone sequential agent transfers. When enable_replay=False it runs in pass-through mode and ctx._workflow_scheduler stays None — the canonical “am I inside a workflow graph” check across ADK.

Key features and the problems they remove

Each advertised feature in ADK corresponds to a piece of plumbing you would otherwise write yourself.

Workflow graph runtime. src/google/adk/workflow/_dynamic_node_scheduler.py and src/google/adk/workflow/utils/_graph_parser.py implement routing, fan-out/fan-in, loops, retry, state, dynamic nodes, and nested workflows. The scheduler is the single driver for both graph-integrated dynamic execution and standalone sequential agent transfers; when replay is disabled it runs in pass-through mode and ctx._workflow_scheduler stays None, which is the canonical “am I inside a workflow graph” check across ADK. This removes hand-written control flow around LLM calls and makes the topology inspectable and replayable from session events.

Task API for agent-to-agent delegation. src/google/adk/workflow/_llm_agent_wrapper.py and src/google/adk/tools/agent_tool.py let a coordinator call a specialist as a tool in single-turn or multi-turn task mode, with the specialist’s output promoted to the task result. Delegation gets a typed contract instead of a free-text handoff.

Tool confirmation / HITL. A tool requiring confirmation emits an adk_request_confirmation function call and pauses the invocation. The response is parsed as JSON if possible, otherwise as a yes/no string, and only a positive value sets confirmed=true. The CLI side lives in src/google/adk/cli/cli.py; tests/unittests/runners/test_run_tool_confirmation.py exercises the flow.

Context caching. src/google/adk/models/gemini_context_cache_manager.py reuses a Gemini explicit cache for the stable prefix and strips the cached system instruction, tools, and contents from the outgoing request. The cache is fingerprint-gated rather than created eagerly, so a prefix that never repeats is never paid for.

Pluggable services. Session, artifact, memory, and credential backends under src/google/adk/sessions/, artifacts/, memory/, and auth/credential_service/ swap between in-memory, SQLite, database, and Vertex AI Agent Engine implementations without changing agent code.

Evaluation harness. src/google/adk/evaluation/agent_evaluator.py and src/google/adk/cli/cli_eval.py run eval sets with tool-trajectory, response-match, safety, and rubric-based metrics, persisted. “Does the agent still work?” becomes a repeatable command.

One caveat: GeminiContextCacheManager, SessionStateCredentialService, TelemetryConfig, and support_cfc are explicitly marked experimental in their docstrings and may change without notice.

Use cases, including where ADK is the wrong tool

ADK fits a multi-agent system where a coordinator delegates to specialists and the whole run pauses for human approval before a destructive tool executes. Task-mode agents, _TaskAgentTool delegation, the adk_request_confirmation flow, and Runner resumability are built in and exercised by tests under tests/unittests/runners/. You are not writing the pause-and-resume machinery yourself.

It also fits local prototyping where you want to inspect the event stream before committing to a deployment story. adk run, adk web, and JSONL event output give a local loop with no infrastructure, and InMemoryRunner needs no services at all. The same holds for deterministic, replayable pipelines that mix LLM steps with plain Python functions and conditional routing: Workflow provides a graph with routing maps, fan-out, loops, retry, and dynamic nodes, and the scheduler supports replay from session events.

It is the wrong tool for a team that wants a minimal, dependency-light agent loop they fully control and can read in an afternoon. The core install pulls in FastAPI, Starlette, OpenTelemetry, and SQLAlchemy-adjacent storage, plus a large optional extra surface, and the workflow and task abstractions add concepts to learn before the first turn runs.

It is a partial fit for a team standardizing on a non-Google model provider and wanting first-class support for that provider’s native features. Non-Gemini models go through adapters (LiteLLM, Anthropic, LangGraph), and several features such as context caching and service tiers are documented as Gemini-only or interactions-API-only.

One flag for anyone evaluating: adk web is documented as development-only and warns it has access to all data and should not be used in production.

Interface and usage: what you actually write

The public surface is a Python library API — Agent, Workflow, Runner, App — plus a Click CLI (adk run, adk web, adk eval, adk create, adk deploy) and a FastAPI dev server. Agents can also be declared in YAML through the Agent Config loader. Start with the smallest thing that runs:

1
2
3
4
5
6
7
from google.adk import Agent

root_agent = Agent(
    name="greeting_agent",
    model="gemini-2.5-flash",
    instruction="You are a helpful assistant. Greet the user warmly.",
)

Agent is a TypeAlias for LlmAgent, so this is the full LLM-backed agent definition. The module-level name root_agent is what the CLI and the eval loader look for — adk run path/to/my_agent imports the module and reads that attribute.

A two-step workflow is the same construction with an edge list:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
from google.adk import Agent, Workflow

generate_fruit_agent = Agent(
    name="generate_fruit_agent",
    instruction="Return the name of a random fruit. Return only the name.",
)

generate_benefit_agent = Agent(
    name="generate_benefit_agent",
    instruction="Tell me a health benefit about the specified fruit.",
)

root_agent = Workflow(
    name="root_agent",
    edges=[("START", generate_fruit_agent, generate_benefit_agent)],
)

The three-element tuple is parsed into two edges, so the fruit agent’s output becomes the benefit agent’s input without any explicit wiring.

Runner executes a turn. Its signature is keyword-only: Runner(*, app=None, app_name=None, agent=None, node=None, plugins=None, artifact_service=None, session_service, memory_service=None, credential_service=None, plugin_close_timeout=5.0, auto_create_session=False). run_async(user_id, session_id, new_message, invocation_id=None) returns an AsyncGenerator[Event, None], so consumption is a plain async loop:

1
2
3
4
async for event in runner.run_async(
    user_id=user_id, session_id=session_id, new_message=content
):
    print(event)

RunConfig carries per-run behavior: streaming_mode, max_llm_calls, tool_thread_pool_config, service_tier, and get_session_config. App binds a root node with app-wide plugins, event compaction, context cache, and resumability config, and is the recommended Runner input.

On the CLI, adk run path/to/my_agent auto-resumes HITL interrupts; adk eval takes an agent directory, an evalset JSON, and a config file path; adk deploy cloud_run|docker|gke --with_ui containerizes and deploys.

One gotcha: GenerateContentConfig fields that duplicate dedicated LlmAgent arguments — system_instruction, response_schema, base_url — are rejected with a ValueError telling you where to move them. Passing them through generate_content_config will not work.

The interesting parts: state merge semantics and two-pass scope resolution

Session state merging uses dict.update semantics, and ADK implements them in SQL rather than reaching for json_patch. The reason is behavioral: json_patch deep-merges nested dicts and treats a null value as a delete-key instruction, while ADK wants a delta key to always win — including when its value is null. From src/google/adk/sessions/sqlite_session_service.py:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
# Merges {delta} into {state} with dict.update() semantics: keys in the delta
# always win with their delta value (including SQL NULL / JSON null), unlike
# json_patch() which deep-merges dict values and treats null as "delete key".
_MERGE_STATE_SQL = """
        SELECT json_group_object(
                 key,
                 CASE
                   WHEN type IN ('object','array') THEN json(value)
                   WHEN type IN ('true','false') THEN json(type)
                   ELSE value
                 END)
        FROM (
          SELECT key, value, type FROM json_each({delta})
          UNION ALL
          SELECT key, value, type FROM json_each({state})
           WHERE key NOT IN (SELECT key FROM json_each({delta}))
        )
      """

The UNION ALL puts every delta key first and appends only the state keys the delta did not mention; json_group_object then reassembles the object. The CASE exists because json_each reports a type per value, and objects, arrays, and booleans must be re-emitted as JSON rather than as strings. The cost is a non-obvious expression that must stay in sync with the Python-side semantics it mirrors.

The second technique is scope resolution for paused task agents in src/google/adk/runners.py. A single backward walk over session events is wrong: it hits post-finish status events before the older success function response and concludes the task is still active. _find_active_task_scope scans forward first to collect finished scopes and aborted invocations, then walks backward to find the newest scope that is neither. That two-pass shape is worth stealing whenever “is this still open?” must be answered from an append-only log.

Two smaller patterns round this out. In src/google/adk/tools/spanner/client.py, the Spanner client is unannotated, so ADK defines _SpannerClient/_SpannerDatabase/_SpannerSnapshot Protocols covering only the operations it uses and casts the real client to them at construction — the SDK’s typing gap stays in one function instead of leaking Any. In src/google/adk/models/gemini_llm_connection.py, Gemini 3.x Live yields tool calls immediately and sends a . placeholder after replayed history because it does not emit turn_complete until it receives the tool response, so buffering would deadlock; 2.5 buffers until turn_complete. Both behaviors sit behind _is_gemini_3_x_live, computed once in __init__ from the model name.

How it compares

This comparison reflects general knowledge and may be out of date. The digest’s confidence for the OpenAI Agents SDK entry is low and for LangGraph is medium; performance is marked unknown for both alternatives, so no cell is filled in.

The key structural point is that ADK adapts LangGraph via LangGraphAgent rather than competing with it at the graph layer. The difference is that ADK also ships the session/artifact/memory/credential services, the CLI, the eval harness, and deployment tooling, at the cost of a much larger dependency surface and a Gemini-first default.

Against the raw google-genai SDK: ADK adds orchestration, state, tool dispatch, HITL, and multi-agent routing on top of the model call; the genai SDK is what ADK itself calls underneath.

Against a hand-rolled asyncio agent loop: ADK’s value is the parts that are tedious to get right — resumability across interrupts, event branch/isolation scoping, session state merge semantics, and toolset lifecycle — not the model call itself.

What to take away

The transferable lessons are mostly about containment. When a system must answer “is this still open?” from an append-only log, scan forward to collect closed scopes before walking backward to find the active one; a single backward pass gets fooled by post-close status events. When a third-party SDK ships without type annotations, define a Protocol covering only the operations you call and cast at the boundary, rather than letting Any spread. When caching a derived artifact, gate creation on a fingerprint match from a prior turn instead of creating eagerly. Version-specific streaming quirks belong behind one boolean computed at construction from the model name. And if you cache clients keyed partly on unhashable inputs, detect that case and skip caching rather than falling back to identity keys.

What ADK does not solve: it is not a hosted platform, not an observability product, and not a model. Several features are marked experimental in docstrings and may change without notice. SessionStateCredentialService stores end-user tokens in session state in clear text. The API moves quickly, and the dependency surface is large.

The repository is at github.com/google/adk-python, Apache-2.0 licensed.