Comparisons
Claude 3.5 Sonnet vs OpenAI o1 for Writing Playwright E2E Tests: Which Engine Best Handles Dynamic DOM Selectors?
Writing flake-free end-to-end tests is a painful chore. We test Claude 3.5 Sonnet and OpenAI o1 on a dynamic, nested React dashboard to see which writes the best Playwright locators.
Updated 10/5/2026
The Nightmare of Dynamic DOM Locators
Writing end-to-end (E2E) tests is the software engineering equivalent of eating your vegetables. We know we have to do it to keep our production pipelines green, but we would rather be doing almost anything else. Dynamic selectors, custom nested shadow trees, asynchronous state changes, and automated build tools that scramble class names (looking at you, Tailwind and styled-components) make writing robust Playwright scripts incredibly tedious.
Naturally, we want to offload this toil to LLMs. But while any basic model can write a script to click a static <button id="submit">, modern single-page applications (SPAs) are a different beast.
We pitted Anthropic’s flagship coder, Claude 3.5 Sonnet, against OpenAI’s reasoning heavy-hitter, OpenAI o1, to see which engine writes clean, resilient, and flake-free Playwright code when faced with a messy, real-world React dashboard.
The Test Scenario
We fed both models an HTML snippet of a highly dynamic dashboard page. The target page features:
1. A sidebar menu that loads asynchronously.
2. A data table where rows are dynamically rendered after an API call.
3. CSS classes generated by a CSS-in-JS library (e.g., class="Button_btn__x92j1 sc-bdVaJa gZgXgQ").
4. A "Delete" button within a specific row that only appears when hovering over that row, triggering a modal confirmation that must be clicked to complete the action.
We asked both models to write a Playwright test script in TypeScript that navigates to the dashboard, locates the row containing the text "Invoice #1042", hovers over it, clicks the hidden "Delete" button, and confirms the deletion in the modal.
---
Claude 3.5 Sonnet: The Semantic Locator Champion
Claude 3.5 Sonnet has built an incredible reputation for developer ergonomics on /platforms/claude. It is fast, intuitive, and usually understands what you mean without requiring a 500-word prompt.
The Approach When presented with the dynamic DOM, Claude immediately avoided the trap of relying on volatile CSS class names. Instead, it leveraged Playwright's modern, locator-first philosophy. It structured the locator using accessibility roles and text content:
`typescript
// Claude's selector strategy
const row = page.locator('tr', { hasText: 'Invoice #1042' });
await row.hover();
const deleteButton = row.getByRole('button', { name: 'Delete' }); await deleteButton.waitFor({ state: 'visible' }); await deleteButton.click();
const confirmButton = page.getByRole('button', { name: 'Confirm Deletion' });
await confirmButton.click();
`
The Pros * **Semantic Locators**: Claude automatically chose `getByRole` and `hasText` filters, which closely mimic how a real human user navigates the DOM. This makes the test highly resilient to future UI redesigns. * **Asynchronous Awareness**: It correctly anticipated that the delete button might not be instantly clickable upon hover, adding a defensive `waitFor` step. * **Readability**: The generated TypeScript was clean, well-commented, and used modern Playwright practices.
---
OpenAI o1: The Overthinking Architect
OpenAI o1 relies on an internal reasoning chain before returning its output. This makes it slower, but in theory, it is supposed to catch the edge cases that fast-token models gloss over. You can read more about OpenAI's reasoning architecture on /platforms/openai.
The Approach o1 spent roughly 18 seconds "thinking" about our DOM. It analysed the nested nature of the components and the possibility of race conditions during the modal transition. It outputted a beautifully structured Page Object Model (POM) pattern:
`typescript
// OpenAI o1's Page Object approach snippet
export class DashboardPage {
private readonly page: Page;
private readonly tableRow = (invoiceId: string) =>
this.page.locator(xpath=//tr[contains(., "${invoiceId}")]);
constructor(page: Page) { this.page = page; }
async deleteInvoice(invoiceId: string) {
const row = this.tableRow(invoiceId);
await row.scrollIntoViewIfNeeded();
await row.hover();
// Defensive check for dynamic element rendering
const deleteBtn = row.locator('button:has-text("Delete")');
await expect(deleteBtn).toBeVisible({ timeout: 5000 });
await deleteBtn.click();
const modalConfirm = this.page.locator('role=dialog >> role=button[name="Confirm Deletion"]');
await modalConfirm.click();
}
}
`
The Pros * **Structural Sophistication**: Rather than outputting a raw script, o1 instinctively organised the code into a scalable Page Object Model, which is the industry standard for production test suites. * **Defensive Design**: It added explicit viewport scrolling (`scrollIntoViewIfNeeded`) to ensure the hover action wouldn't fail on smaller screen configurations. * **Complex Selectors**: The modal interaction locator used a chaining locator (`role=dialog >> role=button[...]`) which is exceptionally robust for handling portals and dynamic modals.
The Cons * **XPath Regression**: It defaulted to using an XPath selector (`xpath=//tr[...]`) for the table row. While highly functional, XPath is generally harder to read and maintain than Playwright's native locator chaining.
---
Which Model Should You Use?
If you want to quickly generate tests directly from your terminal or IDE during a coding session, Claude 3.5 Sonnet is the winner. It writes modern, readable Playwright scripts that prioritise accessibility-first locators, meaning your tests will remain stable even when your CSS frameworks change.
However, if you are setting up a brand-new test suite from scratch and need to architect a robust Page Object Model structure that accounts for tricky browser-level behaviours (scrolling, iframe nesting, viewport constraints), OpenAI o1 is well worth the extra wait time. It acts like a meticulous QA engineer who thinks about screen sizes and layout shifting.
For more hands-on tutorials and troubleshooting tips for your AI-assisted engineering workflows, check out our dedicated platform guides at /platforms/claude/articles and /platforms/openai/articles.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.