Tickd.ai
← The Tickd Guide

Comparisons

Claude 3.5 Sonnet vs GPT-4o for Cypress Test Suite Generation: Which AI Best Handles Dynamic DOM Selectors?

Writing automated frontend tests is the perfect chore to offload to an LLM—until dynamic class names and flaky selectors enter the chat. We pit Claude 3.5 Sonnet against GPT-4o to see which AI actually writes stable Cypress tests.

Updated 10/5/2026

The Flaky Selector Nightmare

Automating Cypress end-to-end (E2E) test suites feels like the ultimate promise of generative AI for developers. In theory, you feed an LLM your component code, ask it for a set of robust integration tests, and watch your build pipeline green-light itself.

In practice, it is rarely that simple. Modern frontends are messy. Tailwind CSS clutters your HTML with hundreds of utility classes, styled-components generate random hashes at build time, and JavaScript frameworks dynamically mount and unmount DOM nodes without warning. An AI that relies on fragile CSS paths like cy.get('div > span.css-1y8b3m > button') will write tests that break the moment you change a margin or redeploy.

To write truly stable tests, an LLM must understand semantic HTML, prioritising durable attributes like data-testid or accessible ARIA roles, while gracefully handling asynchronous state changes. We put the two reigning champions of the coding world—Anthropic's Claude 3.5 Sonnet and OpenAI's GPT-4o—through three rigorous test-writing challenges to see which one builds a suite that actually survives a production deploy.

Round 1: Navigating Tailwind and CSS Module Chaos

For our first test, we presented both models with a complex, interactive React component styled with a mixture of Tailwind utility classes and CSS Modules. The component features a multi-step user onboarding modal that lacks explicit data-testid properties. We instructed both LLMs to generate a Cypress script that clicks through the steps, fills in the forms, and asserts that the final success message is visible.

GPT-4o’s Performance GPT-4o fell back on classic, fragile DOM traversal patterns. It frequently generated selectors targetting specific class chains, such as: ```javascript // GPT-4o's fragile selector strategy cy.get('.flex.items-center.justify-between.mt-4 > button.bg-blue-500').click(); ``` While this technically works in a static local sandbox, it is a ticking time bomb. If a designer changes that button's color to green (`bg-green-500`) or alters the layout spacing, your test suite instantly fails. GPT-4o understands *what* the element is, but fails to anticipate how styling shifts will break E2E testing pipelines.

Claude 3.5 Sonnet’s Performance Claude showed a much deeper appreciation for testing best practices. Instead of clinging to transient CSS classes, it prioritised text-based selection and accessible ARIA attributes, falling back on structure only when necessary: ```javascript // Claude 3.5 Sonnet's robust selector strategy cy.contains('button', /continue/i).should('be.visible').click(); ``` When faced with duplicate buttons, Claude did not default to raw indexes (e.g., `.eq(1)`). Instead, it scoped its assertions using `.within()` blocks, targetting the parent container logically. This is the difference between a test suite that breaks on every commit and one that stands the test of time.

To master this style of robust selector architecture yourself, check out our guide on creating semantic structures using our /prompts.

Round 2: Handling Asynchronous State and Dynamic Loaders

Next, we tested how each model handles the asynchronous reality of modern web apps. We gave them a component that fetches search results from an external API, displays a skeleton loader for an unpredictable amount of time, and then renders the results list. We wanted to see if the LLMs would use modern Cypress waiting practices or resort to anti-patterns.

` +-----------------------------------+-----------------------------------+-----------------------------------+ | Criterion | GPT-4o | Claude 3.5 Sonnet | +-----------------------------------+-----------------------------------+-----------------------------------+ | Asynchronous Strategy | Hardcoded cy.wait(3000) | API Intercepting (cy.intercept)| | Loader Handling | Assumed instant transition | Asserted loader detachment | | Code Flakiness Rating | High | Very Low | +-----------------------------------+-----------------------------------+-----------------------------------+ `

GPT-4o opted for the ultimate cardinal sin of Cypress testing: hardcoded waits. It suggested cy.wait(3000) to allow the API call to resolve. If the server responds in 200ms, you've wasted nearly three seconds of build time; if it takes 3100ms, your test fails.

Claude 3.5 Sonnet instantly set up a network interceptor using cy.intercept(), assigned an alias to the routing endpoint, and used cy.wait('@getSearchResults') to resume execution the millisecond the payload arrived. It also explicitly asserted the disappearance of the skeleton loader (.should('not.exist')) before interacting with the results. This is production-grade code that minimises pipeline runtimes.

Round 3: Token Efficiency and Cost at Scale

If you are using these LLMs via API to auto-generate tests for a massive monorepo of hundreds of components, processing costs and rate limits become highly significant.

  • GPT-4o via the /platforms/openai is highly optimised for raw speed and cost efficiency, coming in at $5.00 per million input tokens and $15.00 per million output tokens. It rarely hits rate limit bottlenecks on standard enterprise tiers.
  • Claude 3.5 Sonnet via the /platforms/claude costs $3.00 per million input tokens and $15.00 per million output tokens. While slightly cheaper on input, Sonnet's output speed can occasionally bottleneck on massive, simultaneous batch runs. If you experience performance degradation or timeout errors, you can consult the official Claude Support Portal to configure asynchronous batch processing.

However, because Claude’s code outputs are significantly cleaner on the first pass, you spend far fewer tokens on iterative debugging prompts compared to GPT-4o, which often requires two or three corrections to remove raw CSS selectors.

The Verdict: Which AI Writes Better Tests?

While GPT-4o is a fast, cost-effective engine for generic scripting, Claude 3.5 Sonnet is the clear victor for automated frontend QA tasks.

Sonnet inherently understands the architectural philosophies of tools like Cypress and Testing Library. It designs tests that mimic real human behaviour (looking for text and accessible labels) rather than brittle machine paths (looking for Tailwind hashes and deep div nests). If you want an E2E test suite that actually saves you development time rather than creating a mountain of maintenance debt, let Claude write your spec files.

codingclaudeopenaicypressqa-automation

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.