# Investigating macOS input-to-display latency

I recently noticed something strange while running typing-latency benchmarks for Koi.

The editor had previously measured very well, but newer measurements were consistently worse.
I initially assumed I had introduced a regression somewhere.

After investigating the editor, Qt, Scintilla, code signing, and various build configurations, I eventually discovered something unexpected.

**The exact same application was significantly faster when launched directly from Apple Terminal than when opened normally through macOS.**

For example, launching the application normally:

```
open -n /Applications/Koi.app
```

Produced higher input-to-display latency than executing its bundled Mach-O directly:

```
/Applications/Koi.app/Contents/MacOS/Koi
```

Same executable, same application bundle, same display, same input source, and same benchmark.

This turned into a much longer investigation than I expected.

I still haven't identified the exact mechanism, but I've reproduced the behavior in several unrelated applications, on two Macs, and across macOS Sequoia and Tahoe.

## First, rule out Koi

My first assumption was that something had changed in Koi's rendering or input handling.

I investigated several possibilities:

- Qt input-method handling
- QScintilla accessibility
- Code signing
- Repaint regions
- Application-side paint times
- Build configuration

Disabling Qt input-method support made no difference.

I also rebuilt QScintilla with accessibility disabled using `QT_NO_ACCESSIBILITY`.
No difference.

Then I instrumented Koi's `paintEvent()` directly.
Paint times were typically well below 1 ms, and both launch methods painted equivalent regions.

Nothing obvious explained the difference.

The benchmark uses [keypress.sh](https://keypress.sh/), which measures actual input-to-visible-pixel latency rather than application event handling.
This distinction probably matters.

An application can process an input event and finish painting quickly, but the corresponding change still needs to pass through the macOS presentation pipeline before it becomes visible.

The additional latency wasn't showing up in my application-side measurements.
So I removed Koi from the experiment entirely.

## A minimal Qt application

I built a minimal application using `QPlainTextEdit`:

```python
import sys

from PyQt6.QtWidgets import QApplication, QPlainTextEdit

app = QApplication(sys.argv)

editor = QPlainTextEdit()
editor.resize(1200, 800)
editor.show()

sys.exit(app.exec())
```

I packaged it as a normal `.app` bundle using py2app.

Then compared:

```
open -n dist/QtLaunchTest.app
```

Against:

```
dist/QtLaunchTest.app/Contents/MacOS/QtLaunchTest
```

Same result.

The normally launched application was slower.

At this point, Koi, Scintilla, custom lexers, and pretty much everything else in my editor were out of the picture.

## The difference follows the display refresh rate

This was the first particularly interesting result.

At 60 Hz, the minimal Qt application measured:

| Launch method | Average | p95 | p99 |
|---|---:|---:|---:|
| Open normally | 29.676 ms | 37.742 ms | 43.080 ms |
| Direct Mach-O | 16.404 ms | 20.500 ms | 24.882 ms |

The p95 difference was **17.242 ms**.
One refresh interval at 60 Hz is 16.667 ms.

I repeated the experiment at 100 Hz:

| Launch method | Average | p95 | p99 |
|---|---:|---:|---:|
| Open normally | 24.687 ms | 31.182 ms | 33.797 ms |
| Direct Mach-O | 16.295 ms | 20.623 ms | 23.286 ms |

Now the p95 difference was **10.559 ms**.
One refresh interval at 100 Hz is 10 ms.

That's remarkably close in both cases.

It suggested that the slower application might be missing a presentation opportunity and reaching the display one refresh interval later.
Not proven, but certainly worth investigating.

## Does this happen in other editors?

I tested Xcode.

On my Mac mini M2 Pro running macOS Sequoia 15.7.5:

| Launch method | Average | p95 | p99 |
|---|---:|---:|---:|
| Open normally | 23.433 ms | 29.886 ms | 31.877 ms |
| Direct Mach-O | 15.787 ms | 19.494 ms | 20.480 ms |

Same behavior.

I then tested several other editors on a MacBook Air M4 running macOS Tahoe 26.6.1.

| Application | Open p95 | Direct p95 |
|---|---:|---:|
| Koi | 39.924 ms | 25.635 ms |
| BBEdit | 38.176 ms | 22.685 ms |
| Xcode | 40.122 ms | 27.233 ms |
| Sublime Text | 56.696 ms | 34.390 ms |

Every application showed lower latency when its bundled executable was launched directly.

Zed also reproduced the behavior on the Mac mini, with p95 increasing from 27.514 ms to 48.208 ms when launched normally.

Different applications, different rendering implementations, different hardware, and different macOS versions.
This was clearly not something specific to Koi.

## Maybe Terminal is doing something?

Initially, I thought the difference was simply between LaunchServices and direct executable launches.
But further testing complicated that explanation.

- Executing the same application directly from Apple Terminal was fast.
- Executing it directly from Ghostty was slow.
- Executing it through an SSH connection to localhost was also slow.

So direct Mach-O execution alone wasn't enough.

Something about the process context appeared to matter.
I tried detaching the application from Terminal:

```
nohup /Applications/Koi.app/Contents/MacOS/Koi </dev/null >/tmp/koi.out 2>/tmp/koi.err & disown
```

Still fast.

I also wrote a small launcher that detached the application and created a new process session.
The resulting application had:

```
PPID = 1
no controlling TTY
session leader
```

Still fast.

So the lower-latency state didn't require Terminal to remain the parent process, an attached terminal, or the same Unix session.
Whatever was responsible appeared to survive those changes.

## Following the process through fork and exec

Next, I tested whether the launch context survived `fork()` and `exec()`.

I started a small launcher from Apple Terminal.
That launcher spawned the Qt test application using fork and exec.

The application remained fast.

```
Apple Terminal
    |
    launcher
    |
    fork / exec
    |
    QtLaunchTest (fast)
```

Then I packaged a launcher as an application bundle and started it normally through LaunchServices.
The launcher replaced itself with the same Qt test executable using `exec()`.

```
LaunchServices
    |
    launcher.app
    |
    exec
    |
    QtLaunchTest (slow)
```

Same final executable.
Different result.

This suggested that the relevant process state was established before the final executable started and survived `exec()`.

## Looking at RunningBoard

I started comparing the process information reported by macOS.

Using `launchctl procinfo`, a normally launched application reported:

```
managed_by = com.apple.runningboard
spawn type = app (1)
spawn role = ui (2)
jetsam priority = 100
started suspended = 1
trampolined = 1
```

It also had an application-specific XPC service:

```
XPC_SERVICE_NAME = application.ai.hackerman.qtlaunchtest...
XPC_FLAGS = 1
```

The fast application launched from Apple Terminal wasn't managed as a RunningBoard application job.
It also reported a different jetsam priority:

```
jetsam priority = 180
```

Interesting.
But these differences didn't establish a cause.

I compared process environment variables, bundle identifiers, QoS, and AppKit activation state.
Changing the bundle identifier manually didn't make the direct executable slow.

Both launch methods reported the same process QoS:

```
qos=0x21
relative=0
```

And the same final AppKit activation state:

```
activationPolicy=0
active=True
```

The Mach task QoS policies were also identical.
There were differences in how macOS managed the processes, but I still didn't know which difference mattered.

## Process coalitions

One particularly interesting difference was process coalition membership.
macOS groups processes into coalitions for resource accounting and management.

When I launched QtLaunchTest from Apple Terminal, it inherited Terminal's coalitions.

For example:

```
Apple Terminal:
    resource coalition: 649
    jetsam coalition:   650

QtLaunchTest:
    resource coalition: 649
    jetsam coalition:   650
```

This was the fast case.

Launching the same executable from Ghostty also inherited its launcher's coalitions:

```
Ghostty:
    resource coalition: 10587
    jetsam coalition:   10915

QtLaunchTest:
    resource coalition: 10587
    jetsam coalition:   10915
```

But this case was slow.

A normal LaunchServices launch received its own application coalitions:

```
QtLaunchTest:
    resource coalition: 10926
    jetsam coalition:   10927
```

Also slow.

The results looked like this:

| Launch context | Coalition inherited from | Result |
|---|---|---|
| Apple Terminal | Terminal | Fast |
| Terminal → fork → exec | Terminal | Fast |
| Ghostty | Ghostty | Slow |
| LaunchServices | Separate application coalition | Slow |

The coalition IDs themselves aren't important.

What's interesting is that coalition membership followed the launch context, and the latency state followed that context too.

I also confirmed the same coalition inheritance behavior with VLC.
Unfortunately, correlation isn't causation.

I attempted to create a new coalition from the fast Terminal context and launch the application into it.

macOS rejected the operation:

```
coalition_create(resource): Operation not permitted
```

So I couldn't isolate coalition membership as an independent variable.
For now, coalitions are an observable difference between the tested process contexts, not an identified explanation.

## Then I discovered something even stranger

While running these benchmarks, I noticed that moving the physical mouse appeared to reduce latency.

Not clicking anything.
Just moving the mouse.

I tested this more systematically.

On the Mac mini at 100 Hz:

| State | Average | p95 | p99 |
|---|---:|---:|---:|
| Mouse moving | 14.656 ms | 18.030 ms | 23.509 ms |
| Mouse idle | 27.446 ms | 36.406 ms | 38.112 ms |

That's an 18.376 ms difference at p95.

I repeated the test at 50 Hz:

| State | Average | p95 | p99 |
|---|---:|---:|---:|
| Mouse moving | 16.614 ms | 20.132 ms | 22.807 ms |
| Mouse idle | 27.827 ms | 34.773 ms | 38.298 ms |

The transition happened quickly.

Start moving the mouse, and the application becomes faster.
Stop moving it, and the higher latency returns.

I reproduced this on the MacBook Air too.

But here's another interesting detail.

**Trackpad movement didn't produce the same effect.**

Neither did synthetic mouse movement or cursor warping.
Only movement from a physical mouse reproduced it in these tests.

## It isn't the visible cursor

I wondered whether macOS was updating the display more aggressively because the mouse cursor was moving.
So I tried separating physical mouse movement from visible cursor movement using:

```
CGAssociateMouseAndMouseCursorPosition(false)
```

This allowed the physical mouse to move while the cursor remained stationary.
The latency improvement remained.
So the effect didn't require visible cursor movement.

I also tried several ways of keeping the application or display system active:

- Continuous Qt timers
- Polling the AppKit event queue
- Continuous repaint requests
- Repeated `NSWindow.displayIfNeeded()`
- Keeping the cursor visible
- Synthetic mouse events
- Scrolling
- `NSProcessInfo` latency-critical activity
- Increased GPU activity

None reproduced the physical mouse result.

At this point, the investigation had moved quite far away from text editors.

## Looking at WindowServer and Core Animation

I used System Trace to compare what macOS was doing during physical mouse movement.

There was substantially more activity in WindowServer, NSEvent, and Core Animation while the mouse was moving.

That seemed consistent with the latency improvement.
But there was a problem with that explanation.

The application launched from Apple Terminal was already fast while the mouse was completely idle.
And in that state, Core Animation activity was relatively sparse.

So increased activity alone wasn't sufficient to explain the lower-latency state.

There appeared to be at least two ways to reach it:

1. Launch the application from a process context that produces the lower-latency behavior.
2. Move a physical mouse while an otherwise slower application is running.

The first persists after launch.
The second changes dynamically.

Both produce a similar reduction in externally measured input-to-display latency.

## Where is the additional latency?

The measurements suggest that the difference occurs somewhere after application-side painting.
The approximate path from keyboard input to visible pixels is:

```
Keyboard input
    |
Application event processing
    |
Document update
    |
Application painting
    |
AppKit / Core Animation
    |
WindowServer
    |
Display presentation
    |
Visible pixel change
```

My application-side instrumentation didn't show the additional delay.
But the external input-to-pixel benchmark did.

And the difference closely followed one display refresh interval.

A possible explanation is that the slower process context causes an update to reach the display one presentation opportunity later.
Conceptually:

```
Fast:
    Input → Paint → Presentation N

Slow:
    Input → Paint → Presentation N+1
```

This would explain the refresh-rate measurements.
But I haven't directly observed the corresponding presentation sequence, so this remains a hypothesis.

The measurements don't yet identify whether the delay comes from Core Animation transaction scheduling, WindowServer, display presentation, or some other part of the pipeline.

## What I've established so far

After quite a few experiments, these are the main findings:

1. The same executable can exhibit substantially different input-to-display latency depending on its launch context.
2. The difference reproduces in Koi, a minimal Qt application, Xcode, BBEdit, Sublime Text, and Zed.
3. It reproduces on macOS Sequoia 15.7.5 and Tahoe 26.6.1, on Apple Silicon Macs.
4. The measured penalty closely follows one display refresh interval at 60 Hz and 100 Hz.
5. Direct executable execution isn't sufficient. Apple Terminal produces the fast state, while Ghostty and SSH do not.
6. The fast state survives fork, exec, terminal detachment, and Unix session changes.
7. RunningBoard state and process coalition membership differ between the tested contexts, but neither has been established as the cause.
8. Application-side paint timing doesn't explain the externally measured difference.
9. Physical mouse movement temporarily reduces latency, even without visible cursor movement.
10. Trackpad movement, synthetic events, continuous repainting, and several other attempts to keep the application active don't reproduce that effect.

The most useful result is probably the minimal Qt reproduction.
It removes nearly all application-specific complexity and makes the behavior relatively straightforward to test.

## Reproducing the issue

The minimal application is just:

```python
import sys

from PyQt6.QtWidgets import QApplication, QPlainTextEdit

app = QApplication(sys.argv)

editor = QPlainTextEdit()
editor.resize(1200, 800)
editor.show()

sys.exit(app.exec())
```

Package it as a macOS application using [py2app](https://py2app.readthedocs.io/en/latest/).

Then compare:

```
open -n dist/QtLaunchTest.app
```

With:

```
dist/QtLaunchTest.app/Contents/MacOS/QtLaunchTest
```

For the second command, launch it from Apple Terminal rather than assuming any terminal emulator will produce the same result.

I used [keypress.sh](https://keypress.sh/) to measure actual input-to-visible-pixel latency.
Each benchmark run contains 200 measurements, reporting average, p95, and p99.

Ideally, repeat the experiment at multiple fixed display refresh rates.

It's also worth comparing the same executable launched from different terminal emulators.

## What's next?

> **I haven't found the exact mechanism yet.**

The most promising areas to investigate are Core Animation transaction timing, WindowServer presentation scheduling, and inherited process state.

In particular, I'd like to identify precisely when an application update is committed and when that update becomes eligible for presentation.
Comparing those timestamps between the fast Terminal launch, slow normal launch, and fast physical-mouse-active state should help narrow things down.

The physical mouse result is especially interesting because it suggests the presentation behavior can change dynamically without restarting the application.
That might provide a more direct way to isolate the mechanism than comparing process launch environments.

For now, the strongest evidence points toward macOS presentation or frame scheduling rather than application rendering.

What started as a suspected Koi performance regression has turned into an investigation of how macOS schedules application updates for display.

And apparently, moving a mouse can make a text editor significantly faster?

