Debugging a UI Test Waiting on the Wrong Process

A UI test that normally finished in under a minute suddenly took almost 14 minutes. The problem was not CI load or a slower authentication flow. XCUITest was waiting for the wrong process to become idle.

Debugging a UI Test Waiting on the Wrong Process

A few weeks ago, one of our nightly UI test runs suddenly became much slower than usual. The suite still passed, which made the issue easy to miss at first, but one test that normally took around 50 seconds ended up taking close to 14 minutes.

Nothing obvious had changed in the flow itself. The test still opened the same hosted authentication screen, entered the same credentials, and eventually reached the application. My first assumption was fairly predictable: because the suite was running tests in parallel, I suspected one simulator or worker had simply ended up competing for CPU.

That explanation did not survive contact with the data.

When I looked at the execution timeline, the slow test had effectively been running by itself. There was no meaningful competing workload that could explain several extra minutes, so I started looking more closely at where the time was actually being spent.

The pattern became obvious pretty quickly. Almost every interaction with the hosted authentication page was followed by roughly a minute of waiting. Tapping the email field, typing, moving to another control, each of those actions completed visually almost immediately, but XCUITest continued waiting long after the UI had responded.

The app itself was not doing sixty seconds of work after each interaction. The test framework was waiting for something else.

The test was waiting on a process we did not control

The authentication flow was displayed through a system web view backed by a process outside the application under test. XCUITest normally waits for the process it is interacting with to become idle before continuing, which is usually useful because it prevents the next assertion from running while animations, transitions, or other UI work are still happening.

The interesting part was which process XCUITest considered responsible for becoming idle.

When the test interacted directly with an element exposed through the hosted web view, the wait became associated with the web view process. That process was not reliably reporting the kind of quiescence XCUITest expected, so the framework would sit there until its wait eventually timed out and move on.

The interaction itself had already succeeded. Most of the extra runtime was simply the framework waiting for confirmation that never arrived.

Once that happened several times during the same login flow, a test that should have taken less than a minute could stretch into something close to 14 minutes without actually doing any additional useful work.

That changed the nature of the problem completely. This was not really a slow authentication flow, and it was not a slow simulator either. It was a synchronization problem between the test framework and a process we did not own.

The fix was much smaller than the investigation

One thing I have noticed about debugging issues like this is that the final code change can be almost embarrassingly small compared with the amount of investigation required to understand why it works.

We still needed the web view element in order to know where on screen to interact. The difference was that instead of performing the tap through the web view element itself, we could derive the screen coordinates from that element and perform the actual interaction through the application under test.

At the UI level, nothing changed. The same point on the screen was tapped and the same control received the interaction.

From XCUITest's perspective, though, the interaction was now anchored to our application instead of the hosted web process. Our application would become idle almost immediately, so the long quiescence waits disappeared.

That was the part I found most interesting. Two operations that looked almost identical from the user's point of view had very different behavior from the automation framework's point of view because they were associated with different processes.

Why retries would have made the situation worse

This kind of problem is exactly where it is tempting to start hardening the test mechanically. Increase the timeout, add another wait, retry the test, or rerun the lane when something looks flaky. Those approaches often sound reasonable because UI tests are noisy and everyone wants the pipeline green again.

In this case, they would have made the situation worse.

A retry of a degraded test could add another ten or more minutes to a job that was already running much longer than expected. Enough retries would push the job closer to its overall timeout and potentially leave us with less diagnostic information instead of more.

Increasing timeouts would have been even less useful. The framework was already waiting for a signal that was unlikely to arrive, so waiting longer would not have made the underlying assumption any more correct.

What helped was being able to see where the time was actually going. Once the execution timeline showed that every interaction with the same process was adding roughly a minute, the investigation became much more focused.

Instead of asking why CI was slow, the better question became why interacting with this particular process was so expensive.

There were actually two different problems

The authentication path had another complication: the quiescence stalls were not the only source of instability. We also had cases where the accessibility bridge used to interact with the hosted web content could become unreliable.

From the outside, both problems looked like “the login test is flaky.” They were not the same failure mode.

That distinction mattered because solving the quiescence issue did not automatically mean the whole authentication path had become reliable. It fixed one specific problem, and we still needed to treat the other one separately.

I have seen this pattern enough times now that I try to be careful about it. Finding one real bug can create a strong temptation to explain every nearby symptom with the same root cause. Sometimes that is correct, but often it is not.

Measure first, then decide whether the fix worked

We also separated measurement from the fix itself. Before relying on the change, we wanted a baseline that made test durations visible in CI, so we had something better than “the suite feels faster now.”

That also forced us to be more careful about what counted as evidence.

If a test has an intermittent failure rate, a handful of consecutive green runs can be encouraging without being especially strong proof that the underlying issue is gone. When everyone just wants a flaky pipeline to stop being flaky, it is easy to overvalue a few successful runs.

Having a baseline made it easier to compare actual behavior before and after the change instead of relying on a short streak of greens.

What stuck with me

The technical details here were specific to XCUITest and a hosted authentication flow, but the broader pattern was more familiar. My first explanation was environmental, while the evidence eventually showed that the interaction itself was succeeding and the test framework was simply waiting on the wrong process.

Once we understood that, the fix became surprisingly small.

A lot of frustrating test problems seem to work this way. The visible symptom happens at one layer, while the real issue belongs to another. The hard part is usually not adding another retry or increasing another timeout. It is figuring out what the system is actually waiting for, and whether it is waiting on the right thing.