Enforcing Animation Performance Budgets in CI

Part of Frame Budgeting & 16ms Targets in Performance Budgeting & GPU Architecture.

The problem

A team fixes a janky menu animation, and three months later it is janky again. Nobody introduced an obvious regression; a shadow was added here, a will-change there, a list grew from ten items to forty. Page-level Lighthouse scores stayed green throughout, because they measure load, not interaction.

Animation regressions need their own budget, measured during the animations themselves.

Root cause analysis: load metrics miss motion

Lighthouse measures a page load. Total Blocking Time and Cumulative Layout Shift are useful and catch some motion problems — a heavy entrance animation inflates TBT, a badly sequenced reveal inflates CLS — but nothing in a default run exercises a menu, a drawer or a filter transition.

Interaction metrics need interactions. Interaction to Next Paint is a field metric; in CI it only exists if the test performs the interaction. That means scripting the same flows that matter to users.

Frames are the right unit. During an interaction, the questions are: how many frames were presented late, how long was the longest, and how much main-thread work happened per frame. The Long Animation Frames API answers the first two from inside the page, with no trace parsing.

Layer count is a leading indicator. GPU memory regressions show up as many promoted layers before they show up as dropped frames on the test machine — and the test machine is usually far faster than a user’s phone.

Noise is the main obstacle. Shared CI runners vary. Budgets set too tightly produce flaky failures and get disabled. Median-of-several-runs, generous thresholds and trend reporting make the check trustworthy.

What to measure, and what each catchesCombine a load check with an interaction check; they catch different regressions.What to measure, and what each catchesCatchesToolNoiseTotal blockingtimeHeavy load-time animation setupLighthouse CIMediumCumulativelayout shiftEntrances that push contentLighthouse CILowLong animationframesLate frames during interactionsIn-page observerLowDropped framesJank during a specificanimationScripted traceHighCompositedlayer countMemory regressionsCDP layer treeVery low
Combine a load check with an interaction check; they catch different regressions.

Step-by-step resolution

Adding an animation budget checkScript the interaction, collect frames, assert a generous threshold.Adding an animation budget check1Pick three or four critical animations, such as menu open and list filter.A focused suite2Write a Playwright test that performs each interaction.Reproducible motion3Install a long-animation-frame observer before the interaction.Frame data from inside the page4Assert no frame exceeds the threshold and total blocking time is bounded.A clear pass or fail5Run each interaction three times and use the median.Less flakiness6Record results per commit so trends are visible.
Script the interaction, collect frames, assert a generous threshold.

Production code pattern

// tests/perf/menu-animation.spec.js
import { test, expect } from '@playwright/test';

async function measureInteraction(page, action) {
  await page.evaluate(() => {
    window.__frames = [];
    if (!PerformanceObserver.supportedEntryTypes?.includes('long-animation-frame')) return;
    new PerformanceObserver((list) => {
      for (const f of list.getEntries()) {
        window.__frames.push({ duration: f.duration, blocking: f.blockingDuration });
      }
    }).observe({ type: 'long-animation-frame' });
  });

  await action();
  await page.waitForTimeout(600);                   // let the animation finish

  return page.evaluate(() => ({
    count: window.__frames.length,
    worst: window.__frames.reduce((m, f) => Math.max(m, f.duration), 0),
    blocking: window.__frames.reduce((s, f) => s + f.blocking, 0),
  }));
}

test('account menu opens without long animation frames', async ({ page }) => {
  await page.goto('/');
  const runs = [];
  for (let i = 0; i < 3; i++) {
    const r = await measureInteraction(page, async () => {
      await page.click('[data-test=avatar]');
      await page.click('[data-test=avatar]');        // close again
    });
    runs.push(r);
  }
  const median = runs.map((r) => r.worst).sort((a, b) => a - b)[1];
  expect(median, 'longest animation frame during menu toggle').toBeLessThan(120);
  expect(Math.min(...runs.map((r) => r.blocking))).toBeLessThan(80);
});

test('filtering the grid does not promote excessive layers', async ({ page }) => {
  await page.goto('/catalogue');
  const client = await page.context().newCDPSession(page);
  await client.send('LayerTree.enable');
  await page.click('[data-test=filter-sale]');
  await page.waitForTimeout(500);
  const { layers } = await new Promise((resolve) => client.once('LayerTree.layerTreeDidChange', resolve));
  expect(layers?.length ?? 0, 'composited layers after filtering').toBeLessThan(40);
});
{
  "ci": {
    "collect": { "url": ["https://staging.example.com/"], "numberOfRuns": 3 },
    "assert": {
      "assertions": {
        "total-blocking-time": ["error", { "maxNumericValue": 300 }],
        "cumulative-layout-shift": ["error", { "maxNumericValue": 0.1 }]
      }
    }
  }
}

Rendering Impact: none in production. These checks run in CI. Their value is catching regressions before they reach devices slower than the CI runner.

The thresholds above are deliberately loose: 120ms for the worst frame during a menu toggle would be a poor experience, but it is far enough from the ~30ms a healthy toggle produces that runner noise will not trip it. Tighten thresholds once the trend data shows the real distribution.

Worst animation frame during menu toggle, by commitMedian of three runs per commit; the budget is 120 ms.Worst animation frame during menu toggle, by commitBaseline28 ms+ backdrop blur74 ms+ 40 menu items96 ms+ shadow animation143 ms
Median of three runs per commit; the budget is 120 ms.

Making failures actionable

A failing budget should tell the author what to look at. Include the worst frame’s attributed script in the failure message by capturing scripts from the long animation frame entries, and attach the trace or a screenshot as a CI artefact. A failure that says “longest animation frame 143ms, attributed to menu.js openMenu, forced layout 61ms” is fixed in minutes; one that says “budget exceeded” is usually disabled instead. The attribution fields are described in debugging jank with the Long Animation Frames API.

Verification checklist

Constraints and trade-offs

  • CI machines are faster than user devices; passing CI does not prove smoothness on a phone.
  • The Long Animation Frames API is Chromium-only, so budgets cover that engine.
  • CPU throttling makes tests more representative but noisier.
  • Layer-count assertions depend on protocol details that can change between browser versions.
  • Every added check costs CI minutes; keep the suite to the animations that matter.

Frequently asked questions

Can Lighthouse catch animation regressions?

Only indirectly. It measures page load, so it catches heavy load-time animation work and layout shifts, but not jank during a menu or filter interaction.

What thresholds should I use for animation budgets?

Start from measured baselines with generous headroom, such as twice the observed median, then tighten as trend data accumulates.

How do I reduce flakiness in performance tests?

Run each interaction several times and use the median, keep thresholds loose, pin the runner class, and disable unrelated background work.

What should a failing animation budget report?

The metric, the measured value, the threshold and the attributed script from the long animation frame entry, plus an artefact such as a trace.