My Thoughts on Graph Engineering

Four frameworks converged on graph execution years before "graph engineering" existed as a term — and their own engineers argued against the diagram.

Originally published elsewhere.

My Thoughts on Graph Engineering — eval is not a leaderboard

What the industry actually adopted, what it quietly rejected, and how to tell which one you need.

On 18 July, Peter Steinberger posted a single line on X: “Are we still talking loops or did we shift to graphs yet?” Within about two days there was a manifesto, a backlash, and something the timeline had started calling a discipline.

Every account I can find says it was a joke. Steinberger was needling the field’s renaming habit. Prompt engineering became context engineering became harness engineering became loop engineering, each rename arriving before the previous one had produced much evidence. A mock eulogy followed and got widely quoted, though the coverage disagrees about who wrote it, which is its own small warning about sourcing anything from this week.

I wrote about loop engineering a few weeks ago. Apparently it died before some of you got round to reading it.

My position

Graph engineering describes something real. The name arrived about two and a half years late, and it names the wrong half of what actually happened.

Here is the short version. Four major frameworks converged on graph execution before the term existed. But read what their own engineers wrote while shipping it and you find OpenAI and Google independently arguing that static declarative graphs get cumbersome, both offering code-native authoring as the fix. The industry adopted graph runtime semantics, meaning durable state, checkpointing, resumption and explicit node boundaries. It moved away from the hand-drawn topology diagram that the July coverage treats as the whole idea.

So the rule I would actually apply is not loop-versus-graph. It is: adopt the runtime, be sceptical of the diagram.

What a graph actually is

Strip the vocabulary and there are three primitives.

Nodes are units of work. In an orchestration graph a node is a function, usually one agent doing one bounded thing, often with its own internal loop.

Edges are transitions. A plain edge is unconditional. When A finishes, run B. A conditional edge is a routing function that reads current state and returns the name of the next node. That is the whole branching mechanism, and it is a finite state machine wearing a new hat. An edge pointing back to a node you have already run is a cycle.

State is a single shared object with a declared schema that every node reads and writes. This is the part people underrate. In ordinary code, data moves between steps as arguments and return values, so the data flow lives in the call stack and exists only while the thing is running. In a graph the state is an explicit typed structure sitting outside any individual node. Nodes do not pass data to each other. They mutate a shared record.

That single choice is where the real properties come from.

Checkpointing and resumption. Because state is external and serialisable, the runtime can snapshot it after every node. Google’s ADK documentation is blunt about what this buys: workflows that pause for a human and resume later, even across process restarts. That is what makes “pause here for approval, resume in three hours” a configuration rather than a rewrite.

Declared parallelism. Fan-out is a property of the topology, not something discovered at runtime. Two nodes with no edge between them can run concurrently. In LangGraph, if both write the same state field you must declare a reducer saying how the writes merge. Other frameworks handle concurrent state differently, so check yours, but the general point holds: a graph forces that decision at design time.

Provenance. Every value in state has a node that wrote it and an edge that led there. Which step produced this claim, and what did it depend on, becomes a query rather than an archaeology exercise.

The distinctive property is not nodes and edges. You can draw those over any system. It is that the topology becomes a data structure rather than a runtime behaviour. In ordinary code the shape of your workflow only exists while it executes and you infer it from traces afterwards. Declared upfront, you can inspect it, validate it, diff two versions, and check every node has a reachable exit before running anything.

That is the trade. You buy inspectability and you pay in rigidity, because anything you did not declare, the system cannot do.

Two different things are being called graphs

This is why a lot of the current coverage contradicts itself.

One sense is the orchestration graph above. Workflow topology, nodes as agents, edges as control flow, shared state.

The other is a knowledge graph. Entities and typed relationships as a retrieval substrate, where an agent traverses connections instead of grepping flat documents.

Both are legitimately graphs. They solve unrelated problems. One is about how work gets sequenced, the other about how information is structured. Several of the competing definitions circulating right now are people arguing past each other about which one they meant. Work out which sense is in play before you evaluate any claim about graph engineering and about half the disagreement evaporates.

This piece is about the orchestration sense.

One word, two unrelated systems: an orchestration graph sequences work through nodes like fetch, draft and judge with control-flow edges and a bounded retry back-edge, while a knowledge graph structures information as entities connected by typed relationships

Where loops actually break

A loop is one agent’s control cycle. Discover, act, verify, repeat until a stop condition holds. The design work is in the exit test. I argued last time that a loop is only as good as its verifier, and I have not changed my mind.

Loops do not fail because tasks get long. A loop handles long tasks fine. They fail when work crosses a boundary the loop cannot represent.

Here is the shape of it from my own work, kept generic. Run an agent that queries an analytics datasource, validates the shape of what comes back, and retries on a bad result. That is a loop and it should stay one. Now add a stage that authenticates against a separate system holding its own credentials and its own failure modes. Then add a third stage that cannot run until a human has signed off on the first stage’s output.

You no longer have a cycle. You have a dependency structure. Retrying the whole thing is now wrong, because the correct recovery point depends on which boundary failed. An auth failure at stage two should not re-run the query at stage one. A rejected sign-off at stage three should send you back to the query, not to the auth. A loop cannot express that difference. It only knows how to go round again.

Those are graph questions: which node produced this artefact, what does the next node require before it can run, which branches can run in parallel, and where do you rewind to when stage three fails for a reason that originated at stage one.

The name is a lagging indicator

Here is the timeline, all of it before 18 July 2026.

LangGraph shipped StateGraph in January 2024 and reached 1.0 in October 2025. Microsoft Agent Framework 1.0 went generally available on 2 April 2026, built around a typed graph workflow that routes data along edges and activates executors when their inputs are ready. Google’s ADK 2.0 reached general availability for Python on 19 May 2026 and for Go on 30 June 2026, with workflows available in Python since March. Google’s own documentation describes the release as transitioning ADK from a hierarchical agent executor to a graph-based execution engine.

Four frameworks, one direction, none of them participating in a naming debate. The last of those shipped eighteen days before the tweet.

The name is a lagging indicator: a timeline from LangGraph in January 2024 through MS Agent Framework 1.0, ADK 2.0 Python and ADK 2.0 Go, ending 18 days before the tweet that started the graph-engineering debate in July 2026

So I was wrong to reach for “nothing changed.” A great deal changed. It changed before the name existed, which makes the tweet a lagging indicator on work already in production.

One correction to the framework lists doing the rounds. Microsoft AutoGen has been in maintenance mode since October 2025, with its last feature release in September 2025. Its README states plainly that it will not receive new features, is community managed, and that new users should start with Microsoft Agent Framework. Several guides published in the last fortnight still list AutoGen as a live option. It has been frozen for nine months.

Here is the part that undercuts the succession framing. A stripped-down retry branch in LangGraph:

def should_retry(state: AgentState) -> str:
    if state["quality_score"] >= 0.85:
        return "end"
    if state["retry_count"] >= 3:  # non-negotiable
        return "end"
    return "retry"

graph.add_conditional_edges(
    "judge",
    should_retry,
    {"end": END, "retry": "retrieve"},  # back-edge
)

That back-edge is a loop, living inside a graph, with a hard iteration cap because an unbounded cycle is the same failure as an unbounded loop and harder to spot on a diagram.

Each node still contains a loop. The graph decides how those loops are bounded and connected. Prompt, loop and graph are three different objects: an instruction, a control cycle, a topology. You do not upgrade from one to the next. You add a layer when the layer below runs out of expressive room.

To be fair to the people writing the guides, almost none of them claim otherwise. The serious ones open by saying most tasks never need this. The succession framing lives in the meme uptake, not in the documentation.

What the vendors actually converged on

This is the finding I did not expect, and it is the reason I think the July framing is off.

Read the vendors’ own reasoning and they are not selling you a diagram.

OpenAI’s practical guide to building agents argues that declarative frameworks requiring you to define every branch and conditional upfront as nodes and edges are good for visual clarity and get cumbersome as workflows turn dynamic, often dragging in a domain-specific language you then have to learn. Their Agents SDK went code-first for that reason.

Google, shipping a graph execution engine, published a post explaining why static graph-based workflows become cumbersome to build and maintain when control flow needs to adapt, and offering dynamic workflows expressed in native Python control flow and asyncio as the answer.

LangChain went the same way. LangGraph’s Functional API, shipped January 2025, lets you write durable workflows as ordinary functions with awaits and get checkpointing and resumption without hand-assembling nodes and edges.

Three organisations, independently, shipping graph execution while arguing against static graph authoring.

That is a coherent position and it is not the one in the July coverage. What won is the runtime: durable external state, checkpoints, resumption, explicit node boundaries, declared merge semantics. What did not win is the picture. The industry decided the graph should be something your code produces, not something you draw first and then fill in.

If you take one thing from this piece, take that distinction. Adopting a graph runtime is a well-supported engineering decision with four production frameworks behind it. Committing to a hand-declared topology is a bet the vendors who build these things have publicly hedged.

What it costs you

Graphs are not free, and the honest version of this argument has to price them.

Orchestration overhead is real. Every node boundary is a serialisation point, a checkpoint write, and a place where state schema drift bites. ADK 2.0’s own migration notes make this concrete: the release added fields to the core event schema, and teams with rigid database columns for session storage get insertion failures until they migrate. That is the tax on durable state, and it lands on your database, not your model bill.

Conversational graph patterns cost more again. A four-agent exchange over five rounds is at minimum twenty model calls, each carrying accumulated history. Fine for offline work where thoroughness beats speed. Expensive for anything latency-sensitive.

There is also a maintenance cost nobody mentions. A declared topology is a second artefact to keep in sync with intent. When requirements change you edit the graph and the nodes, and a stale edge is a bug type checking will not catch.

How I would actually decide

This is the part missing from every version of this argument I have read, including my own first draft of this piece.

Nobody arguing loop versus graph has said how you would test which one was right. The conversation runs on structural intuition. That is the same failure I keep pointing at with benchmarks: a structure that looks correct is not evidence that it performs correctly.

So here is a criterion you can run.

Instrument your current loop and log, per failure, the step at which it went wrong and the step you had to restart from. Then measure the gap. If failures consistently restart where they occurred, you have a loop and it is working. If you are repeatedly discarding correct upstream work to recover from a downstream failure, you are paying a rework tax, and that tax is the number that justifies a graph.

Both numbers come from a loop you already run. Illustrative, not measured. If you cannot produce them, you do not have the evidence to pick a topology.

Second measurement: count how often a run pauses for something outside the system, whether a human approval, an external event, or a rate limit window. If that is rare, checkpointing is overhead. If it is routine, durable state stops being architecture astronautics and becomes the cheapest thing in your stack.

The rework tax: a loop that fails where it restarts costs one step to retry, while a loop that fails late but restarts early discards three steps of correct work — log the failure step, the restart step, and how often a run pauses on something external

OpenAI’s eight months

At DevDay on 6 October 2025, OpenAI launched AgentKit. Its opening complaint about the state of agent development was fragmented tooling: complex orchestration with no versioning, and manual eval pipelines. The two products aimed at exactly those complaints were Agent Builder, a visual canvas for composing and versioning multi-agent workflows with drag-and-drop nodes, and a significantly expanded Evals platform with datasets and trace grading.

On 3 June 2026, both were deprecated on the same day. Agent Builder shuts down on 30 November. Evals goes read-only on 31 October and shuts down on 30 November with it. The published migration path for workflows is the code-first Agents SDK, or Workspace Agents for cases better served by natural language.

I want to be careful about the weight this carries, and more careful than I was in my first draft. Agent Builder never left beta. ChatKit, launched the same day, went generally available and survives. So this is not a company killing a mature product. It is a company withdrawing a beta after eight months and pointing users at the code-first tool its own guide had preferred all along. That is weaker evidence than a GA reversal would be, and it is still evidence, because the direction of the retreat matches what Google and LangChain independently chose.

The detail I would sit with is the other half. The visual orchestration layer and the hosted evaluation layer died in the same notice. Given that the entire argument for graphs is auditability and traceable state, losing the evaluation surface alongside the orchestration surface is the part that should bother anyone taking the auditability claim seriously.

One note on citations

Within days of the term existing, a claim circulated about a $3.1 million Stanford and Anthropic study on graph engineering. Eugeniu Ghelbur went looking and could not find any trace of it. His conclusion is that it was fabricated. I could not locate it either, though two people failing to find something is not proof it does not exist.

Even at that reduced strength the shape is uncomfortable. A conversation whose entire pitch is provenance produced a citation nobody could source, and it spread anyway.

I am not in a position to be smug about this. The first draft of this article contained six factual defects, including a claim that OpenAI had published no reason for a decision they had explained in public, a framework cited as current that had been frozen for nine months, and a launch-to-deprecation interval I had inflated by fifty per cent. Every one of them survived my own reading and died on contact with a primary source. That is the actual failure mode, and it is fast.

What to do on Monday

Stay in a loop when there is one objective, few tools, few branches, no need to resume from the middle, and a human can verify the output quickly. Adding orchestration here buys you a diagram and a new class of bug.

Adopt a graph runtime when the boundaries are real ones: separate owners, separate credentials, separate failure domains, work that genuinely runs in parallel, an approval that blocks, or a rewind target that is not the beginning. Then confirm it with the rework-tax number rather than the intuition.

Be slower about committing to a declared topology than the July coverage suggests. Every vendor shipping this has published reasons to keep the shape in code.

Either way the exit condition is not optional. Bounded retries in a loop, hard iteration caps on back-edges in a graph. The layer above changes what you can express. It does not change what happens when you forget to say when to stop.

The names will keep churning. What has not churned is the verifier, and whether your system can tell you it is wrong and where. That is still the whole job.

Sources

Verified against primary sources:

  • LangGraph StateGraph, initial release January 2024. Functional API (@entrypoint, @task), January 2025.
  • Microsoft Agent Framework 1.0 GA, 2 April 2026. AutoGen README, maintenance mode since October 2025, last feature release September 2025.
  • Google ADK 2.0: Python GA 19 May 2026, Go GA 30 June 2026, Workflow Runtime and graph-based execution engine, adk.dev/2.0 and Google Developers Blog, “Why we built ADK 2.0” and “ADK for Go 2.0”.
  • OpenAI, A practical guide to building agents.
  • OpenAI, Introducing AgentKit, 6 October 2025, including the 3 June 2026 wind-down note. OpenAI deprecations page: Agent Builder and Evals deprecated 3 June 2026, Evals read-only 31 October 2026, both shut down 30 November 2026.

Secondary sources only, not independently confirmed:

  • Steinberger’s post of 18 July 2026 and the reading of it as a joke, via contemporaneous write-ups rather than the original.
  • Eugeniu Ghelbur, The AI Operator, on the unsourced $3.1M study claim.

Samuel McDonnell, July 2026