Comparisons
Claude 3.5 Sonnet vs OpenAI o1 for Architecting Complex State Machines: Which Reasoning Engine Writes Clean XState Logic Without Infinite Lo
We put Claude 3.5 Sonnet and OpenAI's reasoning model o1 to the test, tasking them with building a complex multi-step checkout state machine. Here is which model actually understands deterministic logic.
Updated 10/5/2026
Designing state machines is one of those development tasks where near-perfect isn't good enough. In a complex, multi-step user flow—like an e-commerce checkout with payment retries, split-shipping options, and session timeouts—a single misplaced transition or missing error state results in a frozen UI, lost conversions, and furious users.
Historically, writing robust state machines (such as those using XState) has been a highly manual, brain-melting exercise in mapping out deterministic logic. Naturally, developers have turned to LLMs to generate these configurations. But state charts require deep spatial and logical reasoning.
Today, we are putting Claude 3.5 Sonnet and OpenAI o1 head-to-head. Our goal: generate a complex, production-ready checkout state machine in TypeScript using XState v5, without creating infinite loops, dead states, or messy TypeScript types.
The Challenge: The Multi-Step Checkout Machine
We prompted both models to write an XState v5 machine to govern a checkout flow. The state machine had to handle: 1. Cart Validation: An asynchronous check that can fail if items sell out. 2. Shipping Selection: Dynamic transition based on local vs. international shipping. 3. Payment Processing: An async service call with a maximum of three retries before diverting to a hard failure state. 4. Session Timeout: A global 15-minute timer that transitions the user back to the cart, releasing held inventory.
To make this test as real-world as possible, we required full TypeScript safety, strict context typing, and clean schema definitions.
Round 1: Claude 3.5 Sonnet — Fast, Elegant, But Prone to Edge-Case Blinking
Claude 3.5 Sonnet is a developer favourite for a reason. Its generation speed is blazingly fast, and when combined with Claude’s Artifacts window, it provides a gorgeous, interactive way to inspect code.
How Claude Handled the Logic Claude immediately structured the XState machine using clean, modern XState v5 syntax (which deprecated some v4 patterns like `assign` string syntax). Its code structure was beautifully organized. It correctly typed the context (storing cart items, payment attempts, and shipping methods) and utilised robust TypeScript generics.
Where Claude struggled was the nested state logic of the payment retries. When defining the transition for the payment failure, Claude generated this transition:
`typescript
on: {
PAYMENT_FAILURE: [
{
target: 'processingPayment',
guard: ({ context }) => context.paymentAttempts < 3,
actions: assign({
paymentAttempts: ({ context }) => context.paymentAttempts + 1
})
},
{ target: 'paymentFailed' }
]
}
`
At first glance, this looks clean. However, in XState v5, handling asynchronous services is usually cleaner inside an invoke block using onDone and onError targets. By relying on a manual PAYMENT_FAILURE event, Claude shifted the burden of triggering the failure event onto the UI component rather than letting the machine self-govern the async promise resolution.
Furthermore, Claude missed a crucial edge-case: if the global session timeout triggered exactly as the payment succeeded, there was a potential race condition that could leave inventory locked in a limbo state.
If you find yourself needing to troubleshoot Claude's architectural output for state logic, check out our guide on debugging stateful code at our Claude articles hub.
Round 2: OpenAI o1 — The Deep-Thinking Logical Bulldozer
OpenAI o1 takes a radically different approach. Instead of streaming code immediately, it spends up to a minute "thinking"—generating internal chain-of-thought tokens to map out logical structures before writing a single character of output.
How o1 Handled the Logic During its 45-second thinking phase, o1 systematically mapped out every single permutation of our checkout flow. It realised that the global session timeout needed to act as an outer parent state (a compound state) to cleanly interrupt any active child states (like payment processing or shipping selection) without needing to duplicate the timeout transition on every single node.
Here is how o1 structured the service invocation for the payment step:
`typescript
processingPayment: {
invoke: {
src: 'processPayment',
onDone: {
target: 'success',
actions: 'clearCart'
},
onError: [
{
target: 'processingPayment',
guard: 'canRetryPayment',
actions: 'incrementPaymentAttempts'
},
{
target: 'paymentFailed'
}
]
}
}
`
This is vastly superior. It completely encapsulates the async operation within the machine itself. It correctly separated the actions (incrementPaymentAttempts) from the state configuration, ensuring the logic was decoupled and easily testable. What really impressed us was o1's inclusion of a rollback state: if the checkout failed or timed out, it automatically triggered an action to release held inventory on the backend.
Developer Ergonomics and the Prototyping Loop
While o1 wins on logical architecture, Claude 3.5 Sonnet is still the undisputed champion of the active developer loop.
If you want to quickly build a visual interface to test your newly generated state machine, you can throw Claude's code directly into Figma Weave to generate an interactive frontend prototype. Claude is highly conversational; if you tell it, "Actually, make the shipping step optional for digital products," Claude will modify the code in seconds.
With o1, the feedback loop is slow and expensive. Because o1 charges heavily for output tokens and incurs a high latency penalty for its thinking phase, using it for rapid, conversational prototyping is incredibly inefficient. It is a model you use when you want to "measure twice and cut once."
Pricing, Limits, and API Costs
If you are running these generations programmatically via API, the cost difference is eye-watering:
- Claude 3.5 Sonnet: $3.00 per million input tokens / $15.00 per million output tokens. Rapid, responsive, and incredibly cost-effective for daily engineering workflows.
- OpenAI o1: $15.00 per million input tokens / $60.00 per million output tokens. You are paying a 4x to 5x premium for that logical reasoning pause.
For a single complex architecture draft, running o1 through the web interface is a no-brainer. But for continuous CI/CD automated generations or code-generation agents, Claude 3.5 Sonnet remains the practical baseline.
The Verdict: Which Engine Belongs in Your IDE?
So, which model makes your state architectures tick?
If you are writing highly nested, business-critical logic where an uncaught transition could break a database or lose user data, pay the premium and use OpenAI o1. Its chain-of-thought reasoning excels at identifying edge cases, race conditions, and structural flaws that autoregressive models like Sonnet consistently overlook.
If you are rapidly iterating, prototyping UI flows, or building out standard application state logic, Claude 3.5 Sonnet is the superior daily driver. It writes gorgeous, modern TypeScript faster than you can think, and integrates beautifully with visual prototyping tools.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.