Good LCP, Users Still Waiting: A Startup Performance Case Study from Our Real-Time Voice Web Tool
The short version: fast paint is not the same as ready to work
We recently worked on startup performance for the web voice tool in Owll Translator. Before changing anything, I checked the numbers. Median LCP was 416 ms. By most dashboards, that is green.
Opening the page myself told a different story. The heading and description appeared quickly, but the tool you actually came for, the language pickers, Original and Translation panels, and Start button, took several seconds. In our controlled test, the old median time for the tool to first appear was 5557 ms.
After the change, under the same conditions:
- Median time for the tool to first appear: 5557 ms → 293 ms
- Median LCP: 416 ms → 324 ms
Be precise so this is not misread: the big win is when the tool interface appears, not LCP. LCP moved 416 → 324 ms. That helps, but it is not the story. The 5557 → 293 ms figure comes from a custom “tool visible” probe, a different measurement from LCP.
Also, “good LCP” in the title only describes the median. The three old samples were 6672, 416, and 376 ms. One of those is clearly not good. I will not pretend the old LCP was consistently excellent.
The core question is simple: browser paint metrics measure paint events, while users care about when the task interface appears, whether it responds, and when they can start the task. Related, not the same moment.
On the app side I previously wrote about speech-pipeline latency: Shaving Latency Off Real-Time Speech Translation. That piece is about the path from stopping speech to hearing a translation. This one is only about how the web page starts.
Where the old version got stuck
A render chain wrapped inside an async process
The old client component ran roughly like this:
SSR renders heading and "Preparing" text
→ client loads controller / transport / API adapter
→ create controller
→ bootstrap: language catalog and remaining quota
→ only after bootstrap settles, setController
→ first mount of the real workspace
Simplified:
Promise.all([loadController(), loadTransport(), loadApi()])
.then(([controllerModule, transportModule, api]) => {
const controller = createController(controllerModule, transportModule, api);
controller.bootstrap().finally(() => setController(controller));
});
// No controller: show "Preparing". Controller present: render Workspace.
The first instinct might be that the requests were sequential. They were not. Catalog and quota were already parallel, and this win did not come from parallelizing them.
The real problem: “checks not finished” had been implemented as “main tool not shown.” The checks must exist. What was wrong was hanging the entire workspace render off the end of that async chain. Slow modules, a slow network, or a failed request all delayed the tool.
Why LCP did not raise an alarm
LCP was not tied to the workspace. In the old samples, candidates included both p and img. An early paragraph paint could produce a very low LCP while the tool was still waiting.
I will not generalize this into “LCP always just measures the heading,” and I have no evidence that every LCP came from the heading. On this page, a low LCP had no necessary relationship with users seeing the tool.
Define three kinds of time first
| Time or state | Meaning here | What it proves |
|---|---|---|
| FCP / LCP | Browser Paint Timing / LCP events | Something painted, or the largest candidate in the window painted |
| Tool shell first visible | Custom probe first sees visible layout for Original and Translation headings and the main button | Main tool structure exists; recording is not yet allowed |
| Service ready | Real controller, catalog, available trial quota, and a valid language pair confirmed; Start enabled | User can explicitly start a session; not that a session is connected |
Goal: pull the second far earlier, avoid hurting paint and interaction, and keep the third honest. No faking catalog, quota, or readiness to look “instantly usable.”
The change: full tool first, service checks in the background
New loading flow
Server HTML: full tool + disabled Start
→ first hydration: same workspace + initial snapshot
→ load modules dynamically in the background
→ create and attach controller
→ bootstrap: catalog and quota checks
→ valid snapshot: configure languages, enable actions
→ invalid: keep workspace, show error and explicit retry
The client no longer swaps the whole page based on whether a controller exists. It always returns the same workspace with a nullable controller and an initialization state. Once created, the controller is attached, then we wait for bootstrap:
async function initialize() {
if (!ownedController) {
const modules = await loadOptionalVoiceModules();
ownedController = createController(modules);
setController(ownedController);
}
await ownedController.bootstrap();
setInitializationState('ready');
}
The first screen shows the same language controls, panels, Start, and Settings. During loading, only status text, confirmed language names, and enabled states change. No non-interactive image as a stand-in, no fake languages or balance.
A common misconception: in Next.js, use client declares a client boundary. It does not mean there is no server HTML. With JavaScript off we still saw real controls and transcript panels, so the initial workspace is in the HTML.
A stable initial snapshot keeps hydration consistent
The workspace uses useSyncExternalStore. Since the controller can be null on first render, we need a consistent, frozen, stable initial snapshot:
const INITIAL_VOICE_SNAPSHOT = Object.freeze({
phase: 'idle',
configuration: Object.freeze({
mode: 'manual', sourceLanguage: null, targetLanguage: null,
}),
activeRunId: null,
languages: null,
transcripts: Object.freeze([]),
remainingSeconds: null,
trialAvailability: 'unknown',
backendAvailability: 'unavailable',
audioBlocked: false,
autoPlayEnabled: true,
error: null,
});
These values are deliberate: languages: null means the catalog is unconfirmed (not “exactly two languages”); remainingSeconds: null means quota is unknown (not “120 seconds left”); phase: 'idle' means no session yet, not that the service is available. Status text is “Checking voice service,” and Start stays disabled.
const getServerSnapshot = () => INITIAL_VOICE_SNAPSHOT; // outside the component
const subscribe = useCallback(
(listener) => controller ? controller.subscribe(listener) : () => undefined,
[controller],
);
const getSnapshot = useCallback(
() => controller?.getSnapshot() ?? INITIAL_VOICE_SNAPSHOT,
[controller],
);
const snapshot = useSyncExternalStore(subscribe, getSnapshot, getServerSnapshot);
Trap: return the same module-level object from getServerSnapshot, do not allocate a new one every read. Server and first hydration share content; after attach, subscribe to real state. A stable reference also avoids unnecessary updates.
Browser tests store DOM references to the H1, workspace, language buttons, panels, status node, and Start, then compare them after bootstrap to confirm main nodes were not replaced.
Showing the tool early is not allowing recording early
Start enables only when initialization finished, a real controller exists, backend and trial status are available, the catalog is non-empty, quota is confirmed and above zero, source and target both exist and differ, and there is no “session creation result unknown” state.
While loading, users can open Settings. Language selection, Start, and session-affecting settings stay gated. Microphone and session creation run only after Start.
Easy bug: controller.bootstrap() may catch an error, write it into the snapshot, and still resolve. So initializationState = 'ready' only means the flow finished, not that the backend is available. The workspace still checks backend, trial, and error fields for error and retry UI. Treating resolve as business success is a classic failure mode here.
Do not turn render waiting into input lag
No new dependencies, no SDK rewrite. Three constraints.
Keep the dynamic module boundary
Controller, transport, and API adapter still load via dynamic import in the background. Nothing was pulled into synchronous first-screen execution or into the Start click handler. The workspace does not wait for those imports to show.
Honest caveat: dynamic loading is not zero download, parse, or main-thread cost. Median total long-task time went from 182 ms to 191 ms, up 9 ms. “Async import” is not “SDK overhead removed.”
A closed language picker does no sorting or rendering
The catalog has 136 language and region options. Closed, the picker returns a stable empty array. Sorting, filtering, and mounting happen only when it opens:
const sortedLanguages = useMemo(
() => pickerField ? [...languages].sort(compareLanguageLabels) : EMPTY_LANGUAGES,
[languages, pickerField],
);
const matchingLanguages = useMemo(() => {
if (!pickerField) return EMPTY_LANGUAGES;
const query = languageSearch.trim().toLocaleLowerCase('en');
return sortedLanguages.filter((language) => !query ||
`${language.label} ${language.code}`.toLocaleLowerCase('en').includes(query));
}, [languageSearch, pickerField, sortedLanguages]);
That avoids option DOM while closed and avoids reprocessing the catalog on every snapshot update. Sorting moves to the first open, so measure it; do not claim INP must improve from this alone. No virtual list, Web Worker, or new state library.
Prefer native interactions
Settings uses <details> / <summary>. Language selection uses <dialog> and a native list. Settings still expands with JavaScript off. That describes this case; it is not a claim that native controls meet every accessibility requirement on every device.
Keep control sizes stable
Getting the workspace onto the first screen is not enough. On phones, “Choose language” wraps differently from “English (United States)” or “Chinese (Simplified),” so arriving data can push transcript panels down.
We reserved the name area with min-h-[3.75rem] and line-clamp-3 (60 px in the test environment). Full names stay in the title and picker. An early two-line reservation was too small; three lines was a real fix, not shortening names to hide the jump.
A controlled slow-init test at 360×780 and 1280×900 compared x/y/width/height for source and target language values, transcript area, and both panels. Before and after service confirmation, every difference was at most 1 px; language value height stayed 60 → 60 px. Across three mobile samples, Original/Translation headings and Start had a y difference of 0 from first appearance to service ready.
Design failure and retry with the first screen
- Module loading has a 20-second bounded wait; unmount releases the wait and clears timers. Existing ~10-second bootstrap and API timeouts remain. Native dynamic import downloads cannot truly be cancelled by AbortSignal; what cancels is waiting and later use of the instance or state.
- Retry reuses the controller and stays single-flight via
initializationPromise. A double-click will not create two controllers or fire two bootstrap rounds. - Retry only re-checks the service. It does not auto-start recording, switch guest identity, restore exhausted quota, or re-send create for a session whose creation result is unknown.
- On hidden, pagehide, and unmount, controllers still stop or dispose. Media cleanup was not removed for a faster first screen.
Forced catalog and quota failures plus double-click Retry: each endpoint saw one additional bootstrap request, identity unchanged, no microphone, create, or poll.
How we measured
Setup
Two exact production builds, run locally, with code version, build ID, and component content hash verified.
| Condition | Setting |
|---|---|
| Browser | Chromium 153 |
| Node | 20 |
| Build | Production (no dev/HMR) |
| Mobile viewport | 360×780, deviceScaleFactor=1 |
| CPU | CDP throttle 4× |
| Samples | 3 fresh browser contexts per version |
| APIs | Same 136-item catalog and 120-second quota fixtures |
| API latency | Fixed 5000 ms per POST, checks still parallel |
| Session | No Start, mic, room, or poll |
| Interaction | 20 Settings toggles; open source picker; type Japanese; clear; Escape |
Limits: a fresh context isolates browser state, not OS, server, or every cache. Three local runs are not all devices and networks. 4× CPU is not a specific phone. This comparison used neither a real backend nor a real microphone, unlike the app-side pipeline article.
”Tool visible” probe
requestAnimationFrame checks Original heading, Translation heading, and the main button: exist, non-zero rect, not hidden or transparent. First satisfaction records performance.now():
function hasVisibleLayout(element) {
if (!element) return false;
const rect = element.getBoundingClientRect();
const style = getComputedStyle(element);
return rect.width > 0 && rect.height > 0
&& style.display !== 'none'
&& style.visibility !== 'hidden'
&& Number(style.opacity) > 0;
}
This is a visible-layout DOM probe, not standard paint timing. It does not fully check viewport intersection, occlusion, or presented pixels. So 293 ms is not “pixel-perfect painted” and not “Start available.” Screenshots, JS-off HTML, and DOM retention/geometry assertions back it up.
Other metric definitions
- LCP: last candidate before first input; local controlled window, not CrUX.
- Layout shift: exclude hadRecentInput, sum before first input; not standard CLS session windows.
- Event Timing: duration ≥ 16 ms, grouped by interactionId; not standard INP.
- Playwright Settings click timing includes auto-wait overhead.
Results
| Metric | Old | New | Notes |
|---|---|---|---|
| Tool visible probe, median | 5557 ms | 293 ms | Main win, ~5.26 s earlier |
| FCP, median | 416 ms | 324 ms | Paint metric |
| Pre-input LCP, median | 416 ms | 324 ms | Do not swap with tool time |
| Pre-input LCP, sample max | 6672 ms | 772 ms | Max of three, not field p95 |
| Pre-input layout shift sum, median | 0.293568 | 0.000218 | Lab proxy, not standard CLS |
| Pre-input long task total, median | 182 ms | 191 ms | 9 ms worse |
| Settings script click p98 | 79 ms | 82 ms | 3 ms worse, includes Playwright |
| Event Timing raw event p98 | 120 ms | 72 ms | Event level, not INP |
| interactionId group p98 | 64 ms | 64 ms | Controlled interaction proxy |
Three samples:
| Version | Tool probe ms | FCP ms | Pre-input LCP ms and candidate |
|---|---|---|---|
| Old | 6333 / 5557 / 5503 | 1188 / 416 / 376 | 6672 img / 416 p / 376 p |
| New | 345 / 293 / 266 | 376 / 324 / 296 | 772 img / 324 p / 296 p |
Every old tool mark landed after the fixture API responses; every new mark landed before them. Service confirmation still completes after the 5-second API delay. We did not make a 5-second backend call return in 0.29 seconds. We changed when users see the tool.
In this controlled interaction set, interactionId-grouped p98 stayed 64 ms. We do not have enough real-user INP data to claim INP is unchanged everywhere.
One new-version desktop sample: tool probe ~310 ms, FCP 344 ms, LCP 756 ms, layout shift sum 0. No desktop baseline, so observation only.
Functional checks: JS off, HTML still has real language buttons, Original/Translation, disabled Start, and “Checking voice service.” With both APIs held 5 seconds, workspace complete, Settings responsive, languages and Start disabled. Node 20 typecheck, production builds, 74 voice unit tests, and related browser regressions passed.
Takeaways
- Show the real workspace early; keep unknown capabilities disabled. Users see structure and what is being checked.
- One workspace plus one stable snapshot beats separate loading and ready screens. Two screens drift.
- Keep dynamic loading, but do not claim it removed cost. Long tasks +9 ms stays in the report.
- Lazy work moves cost. Closed pickers skip sorting; the first open pays, so test with real clicks and typing.
- Check layout stability for dynamic data separately. One wrapped language name can push a panel.
- A resolved Promise is not business success. Failure snapshots must drive clear retry UI.
- Identical conditions; keep outliers and regressions. An average cannot replace task experience.
Still ahead: standard web-vitals with field INP/CLS, more device and network samples, a stricter viewport/paint tool-visible metric, and a service-ready business metric. No full RUM setup yet.
LCP measures “what got painted,” not “can the user start working.” For a tool page, putting the real workspace on screen first and keeping service confirmation in an explicit background state can take time-to-tool from over 5 seconds to roughly 0.3 seconds, if you also protect interaction, layout, and correctness.
Try the current startup feel at Owll Translator.