September 26, 202617 min read

App Optimization for Engineers: Cut P99 Cold Starts and Crashes

App Optimization for Engineers: Cut P99 Cold Starts and Crashes ! Engineer testing mobile app startup performance App optimization is the ongoing engineering work of reducing crashes, ANRs, and tail latency (P95/P99) so users get fast, reliable interactions every time they open your app.

Usama Ahmed Memon
Co-Founder at Bitrupt
App Optimization for Engineers: Cut P99 Cold Starts and Crashes
Engineer testing mobile app startup performance

App optimization is the ongoing engineering work of reducing crashes, ANRs, and tail latency (P95/P99) so users get fast, reliable interactions every time they open your app. Start by fixing crashes and shrinking cold startup time before chasing smaller wins. Tools like Baseline Profiles and Android’s platform-level profiling should be your first stops, not your last resort.

TL;DR:
  • Focusing on tail latency metrics like P95 and P99 is crucial, as averages can hide significant delays affecting a small but important percentage of users.
  • Implementing Baseline Profiles and lazy initialization can improve cold startup times by up to 30%, especially on first app launch.
  • Diagnosing jank requires FrameTimingMetric or trace tools like Perfetto to identify specific causes such as deep layouts or synchronous resource loads.
  • Continuous measurement of crash, ANR, frame drops, and resource usage, segmented by device and OS, is essential to catch regressions before they impact retention.
  • Prioritize fixing critical performance issues identified via production data and automated benchmarking before expanding to tool acquisitions or additional optimizations.

BitruptBuild A Faster, More Reliable AppBitrupt provides custom software development with senior engineers for scalable, secure platforms and ongoing support.Explore Bitrupt

Table of Contents

What app optimization covers and why it affects retention

Think of app optimization the way you’d think about a car’s tune-up. You don’t just check the oil once and call it done. You track engine temperature, tire wear, and fuel efficiency over time, because small drops in performance compound into big problems. The same logic applies to your app: startup time, latency, stability, memory use, and battery drain all need continuous measurement, not a one-time audit.

The stakes are real. Android vitals set bad-behavior thresholds that directly affect your app’s visibility on the Play Store, including an overall user-perceived crash rate of 1.09% and an ANR rate of 0.47%. Cross those lines and Google may reduce your app’s discoverability, regardless of how polished your UI looks.

Here’s the part many teams get wrong: averages hide the problem. A median load time of 800 milliseconds sounds fine until you realize your slowest 5% of users are waiting four seconds or more. Those users are the ones leaving one-star reviews and uninstalling.

What to measure instead of relying on medians:

  • Startup timing at P50, P95, and P99, not just the average.
  • Frame rendering broken down by device tier, since a budget Android phone from 2021 behaves nothing like a current flagship.
  • Crash and ANR rates segmented by OS version, since a regression on one Android release can hide inside a healthy overall number.
  • Battery and memory usage, which quietly erode retention even when nothing technically crashes.

Optimizing for the middle of the curve while ignoring the tail is like designing a bridge that holds up fine on a calm day and collapses the first time it’s actually tested.

Android startup best practices: Baseline Profiles and lazy initialization

Cold startup is the first impression your app makes, and it’s also one of the most fixable problems in mobile engineering. When a user taps your icon, the Android runtime has to load and interpret your app’s code before anything appears on screen, and without help, a lot of that code runs through the slower interpreter or JIT compiler path on first launch.

Baseline Profiles solve this through profile-guided optimization (PGO). According to Android Developers, Baseline Profiles can improve code execution speed by roughly 30% from the very first launch, because critical code paths get ahead-of-time (AOT) compiled instead of waiting on the JIT to warm up. That’s a meaningful head start for an app that only gets one chance at a first impression.

Here’s how to put this into practice:

  1. Generate a Baseline Profile using the Macrobenchmark library. The Android Baseline Profiles codelab walks through writing a benchmark that captures your app’s most common user journeys, like opening the app and scrolling a feed.
  2. Run the profile generator against an unminified build. Method signatures need to match exactly, so profile generation happens before obfuscation, while your release build variant keeps R8 minification for the actual shipped APK.
  3. Verify the improvement with Macrobenchmark. In one sample measured through the codelab, teams saw a substantial improvement in time to full display after applying the profile, a difference you can confirm yourself by comparing compilation modes (None, Partial, Full).
  4. Layer in startup profiles for DEX layout. Startup profiles reorder your app’s DEX files so startup-critical classes load first, complementing what Baseline Profiles already do for method execution.
  5. Delay non-critical initialization with the App Startup library. Anything that doesn’t need to run before the first frame, analytics setup, ad SDKs, background sync, should initialize lazily instead of blocking the launch path.

Pro Tip: Re-generate your Baseline Profile every time you make a major change to a critical user journey. A profile built against last quarter’s onboarding flow won’t help this quarter’s redesigned one.

The metrics that matter here are time to first frame (when something visible appears), time to full display (when the screen is actually usable), and the percentile breakdown of both. A 200 millisecond median startup time means little if your P99 sits at 3 seconds on older devices, and that’s exactly the kind of gap Baseline Profiles are built to close.

Rendering and jank: frame budgets and how to fix slow frames

Every frame your app draws has a deadline. At 60Hz, you have about 16 milliseconds to produce a frame before the system has to show a stale one, a skipped frame the user perceives as a stutter. Higher refresh rate screens shrink that window further: roughly 11 milliseconds at 90Hz and about 8 milliseconds at 120Hz. Miss the deadline often enough and your app feels janky even if nothing ever technically crashes.

Diagnosing the problem starts with the right traces. Use FrameTimingMetric inside Macrobenchmark to capture frame duration data automatically, or dig into raw traces with Systrace or Perfetto for a frame-by-frame view of what’s eating your budget. A single long-running measure pass, an unexpected bitmap decode, or a synchronous database read on the UI thread can each single-handedly blow past 16 milliseconds.

Common fixes worth trying first:

  • Flatten deep layout hierarchies, since every nested ViewGroup adds measure and layout passes.
  • Move bitmap decoding and uploads off the main thread, because GPU texture uploads triggered synchronously are a classic jank source.
  • Enable RecyclerView prefetch and reduce the number of distinct view types, which cuts down on layout inflation during fast scrolling.
  • Avoid allocating objects inside onBindViewHolder or onDraw, since repeated allocations trigger garbage collection pauses at the worst possible moment.

One in twenty of your users experiencing dropped frames on a specific device model is the kind of signal a P95/P99 view catches and an average completely hides, which is why measuring performance guidance from Android Developers recommends frame metrics at P50, P90, P95, and P99 rather than a single mean value.

To catch regressions before they ship, write instrumentation tests that assert on frame timing thresholds using Macrobenchmark, and run them in CI against a fixed device profile so a slow build doesn’t quietly ship a stutter.

Rendering and jank: frame budgets and how to fix slow frames — overview diagram

Memory, allocations, and leaks: finding hotspots before they cause crashes

Memory pressure is a slow leak, literally and figuratively. An app that gradually consumes more RAM during normal use eventually forces the operating system to kill background processes, including your own app when the user switches away and back, which shows up in your metrics as a mysteriously slow resume time or an outright relaunch.

Two tools do most of the diagnostic work: the Android Studio Profiler’s Memory Profiler for live heap inspection, and Simpleperf for sampling CPU and allocation activity at a lower level. Pull a heap dump when memory climbs unexpectedly and look for objects that should have been garbage collected but weren’t, usually a sign of a lingering reference, an unregistered listener, or a static field holding onto a context it shouldn’t.

Patterns that reliably reduce memory pressure:

  • Reuse objects on hot paths instead of allocating fresh instances inside loops or scroll callbacks.
  • Recycle bitmaps explicitly when working with large images, rather than trusting garbage collection timing.
  • Use RecycledViewPool when nesting RecyclerViews, so child views get reused across parent items instead of rebuilt from scratch.
  • Audit listener registration, since forgetting to unregister a broadcast receiver or observer is one of the most common leak sources in production apps.

Validate improvements by comparing two numbers before and after a fix: the time between garbage collection events (longer is better, since it means less allocation pressure) and peak memory usage during a representative user session. If both move in the right direction, you’ve fixed a real problem rather than a symptom.

Network and background work: cutting unnecessary resource use

Startup is not the time to make network calls you don’t strictly need. Every request your app fires before the first frame renders adds latency the user has to sit through, so defer anything that isn’t required to show the first meaningful screen.

Background work deserves the same discipline. On Android, WorkManager lets you batch and delay tasks like sync jobs or log uploads until conditions are favorable, such as when the device is charging or connected to Wi-Fi, instead of firing them the moment a trigger occurs. On iOS, assigning the correct Quality of Service (QoS) level to background tasks keeps non-urgent work from competing with user-visible operations for CPU time, a point Apple’s WWDC guidance emphasizes directly: prioritize user-visible work and keep it off tasks that don’t need to run immediately.

A few habits keep resource use in check:

  • Batch network and disk writes rather than firing them one at a time as events occur.
  • Limit wake locks and background scans, since Android vitals flags excessive partial wake locks above 5% of sessions as a bad-behavior signal that hurts visibility.
  • Watch excessive battery usage, which Android vitals sets at a 1% threshold before it starts counting against your app.
  • Slice your telemetry by device and network type, since a weak-signal user on a three-year-old phone experiences your background sync very differently than someone on Wi-Fi with a current flagship.

Pro Tip: Test your background sync logic on a throttled network profile in Android Studio before shipping. What looks instant on office Wi-Fi can time out entirely on a train.

Monitoring and tooling: what to track and how to catch regressions

You can’t optimize what you don’t measure, and you can’t trust what you only measure once. A production monitoring stack needs to track crash rate, ANR/hang rate, startup timings, frame percentiles (P50, P95, P99), battery consumption, and disk write volume, continuously, not just during a pre-launch QA pass.

Here’s how the tooling maps to each job:

  1. Macrobenchmark for controlled, repeatable performance tests you can run locally and in CI.
  2. Perfetto for deep trace analysis when you need to understand exactly what happened during a slow frame or a startup delay.
  3. Android Studio Profiler for live inspection of CPU, memory, and network activity during development.
  4. Xcode Instruments for the equivalent iOS-side profiling, covering time, allocations, and energy.
  5. MetricKit for aggregated, privacy-preserving performance data collected directly from real user devices on iOS.
  6. Firebase Performance Monitoring for automatic collection of startup metrics and network traces in production, across both platforms, with custom traces for anything Firebase doesn’t capture by default.

The workflow that ties this together runs macrobenchmarks as part of your CI pipeline, not just on demand, and treats any regression in the release pane as a blocker worth investigating before merge. Set alerts on P99 specifically, sliced by device model, since a regression that only shows up on one chipset will disappear entirely from an aggregate view.

When you’re deciding what to fix first, a simple prioritization model works well: impact per user multiplied by the number of affected users — an approach outlined in conversion rate optimization tips that links performance improvements to retention and conversion gains. A crash that hits 0.5% of sessions on your most popular device model usually outranks a jank issue affecting 0.1% of sessions on a rare one, even though both look small in isolation.

Optimization workflow and sprint checklist

Turning analysis into shipped fixes takes a repeatable process, not a one-off heroics sprint. Here’s a workflow that fits into a normal two-week cycle:

  1. Triage using production data. Pull your top crash and ANR clusters and your worst-performing P99 user journeys from the last release.
  2. Reproduce and benchmark. Write a Macrobenchmark test or capture a Perfetto trace that reliably reproduces the issue on a representative device.
  3. Apply a targeted fix. Resist the urge to refactor everything nearby, since a fix that’s hard to isolate is a fix that’s hard to verify.
  4. Validate on an obfuscated release build. A fix that works in debug but wasn’t tested against R8 minification can quietly break in production.
  5. Verify in CI and monitor post-release. Confirm the benchmark passes in your pipeline, then watch the relevant metric in production for at least one full release cycle before declaring victory.

A sample set of service-level objectives (SLOs) and the tickets they generate:

MetricTarget SLOTicket trigger
Crash rateBelow Android vitals’ 1.09% thresholdAny release exceeding threshold for multiple days
ANR rateBelow Android vitals’ 0.47% thresholdAny spike above threshold on a specific device cluster
Cold startup P95Under 2 seconds on mid-tier devicesRegression detected in CI macrobenchmark
Frame drop rateUnder 1% of frames exceeding 16ms budgetJank flagged in FrameTimingMetric trace

This kind of measure-first discipline is the same argument we’ve made before about skipping analysis in app development more broadly: shortcuts feel faster in the moment and cost more later.

Common optimization mistakes and quick ways to avoid them

Most performance problems aren’t caused by lack of effort. They’re caused by measuring the wrong thing or fixing without proof.

  • Chasing median performance instead of the tail. Fix this by setting your SLOs on P95/P99 rather than the average, since that’s where real user pain lives.
  • Optimizing without a reproducible benchmark. A fix you can’t measure is a guess dressed up as engineering; write the Macrobenchmark test first.
  • Blocking the main thread with I/O or heavy allocations. This is still the single most common cause of jank and ANRs, years after tooling made it easy to catch.
  • Skipping obfuscated build testing for Baseline Profiles. A profile generated against the wrong method signatures silently fails to apply in production.
  • Ignoring device-model and OS-segmented telemetry. An aggregate crash rate of 0.8% can hide a 4% crash rate on one popular device, and that’s the number that actually matters to those users.

Bitrupt’s engineering approach to app optimization

This guide reflects the same discipline Bitrupt applies inside client engagements: measure before touching code, and treat obfuscated release builds as the real test, not debug builds. Bitrupt staffs projects with senior engineers only, working through development pods or direct staff augmentation, so performance work gets tackled by people who’ve profiled production apps before, not learned the tooling on the client’s clock. That model shows up in how Bitrupt’s engineering teams operationalize the checklist above: benchmark first, fix second, verify in CI third.

Three priorities to act on this week

Pick one leading metric, either crash/ANR rate or P99 cold startup, and commit to moving it. Then: triage your worst clusters, generate a Baseline Profile, and wire a macrobenchmark into CI. Everything else can wait a sprint.

Optimization techniques that differ between iOS and Android

Android and iOS share the same underlying goals, fast startup, smooth frames, low memory pressure, but the tools and levers differ enough that a one-size playbook doesn’t work.

On Android, Baseline Profiles and startup profiles are the biggest lever available, since they directly change how the runtime compiles and loads your code on first launch. Android also gives you finer control over background work scheduling through WorkManager, and Android vitals provides explicit, published thresholds for crash rate, ANR rate, and battery usage that tie directly to Play Store visibility.

On iOS, there’s no equivalent AOT compilation step to tune, so the emphasis shifts toward minimizing main-thread work during the launch window itself. Apple’s own WWDC guidance frames app launch as a high-intensity period where every synchronous task competes for a narrow window, and recommends offloading non-UI work with the correct Quality of Service level rather than letting it run inline. MetricKit then becomes your main field-data source, since Apple doesn’t expose the same granular vitals dashboard Android developers get through Play Console.

The practical takeaway: Android rewards proactive compilation tuning before launch even happens, while iOS rewards disciplined thread management during the launch window itself. Teams building on both platforms need two distinct playbooks, not one translated version, and Bitrupt’s iOS development guidance covers the platform-specific groundwork for teams starting from scratch.

Android and iOS optimization comparison

How optimization changes what the user actually feels

Performance work is invisible when it’s done right, and that’s the point. A user doesn’t consciously register “cold startup dropped from 2.1 seconds to 1.4 seconds.” They just notice the app feels responsive, or they don’t notice anything at all, which is exactly the outcome you want.

Perceived performance matters as much as raw numbers. An app that shows a skeleton screen or a progress indicator within 100 milliseconds feels faster than one that shows nothing for a full second, even if the actual data takes the same total time to load. This is why time to first frame matters separately from time to full display: users judge responsiveness by what appears on screen, not by when the underlying work technically finishes.

The retention math follows directly from this. Users who hit a crash, a frozen frame, or a slow launch don’t file a bug report, they just leave, often without ever telling you why. That’s the quiet cost behind Android vitals’s thresholds: they’re not arbitrary engineering targets, they’re a proxy for the churn that happens below the surface, and closing the gap between P50 and P99 performance is often the highest-leverage retention work available to a team, well before adding a new feature would move the needle.

Why the tail matters more than the toolkit

The conventional advice on app optimization treats it like a shopping list: add Baseline Profiles, add a profiler, add a monitoring SDK, done. That misses the actual discipline. Tools don’t fix anything by themselves; they just make the tail latency visible, and most teams still aren’t looking at it. I’d argue the single most underrated habit in this entire field is simply setting an SLO on P95/P99 and refusing to ship a release that regresses it, before adding a single new tool to the stack.

What’s overrated: chasing a marginal median improvement on a metric that’s already healthy. What’s underrated: obfuscated-build testing, since so many Baseline Profile failures in production trace back to a profile generated against the wrong method signatures. If you take one thing from this guide, prioritize measurement discipline over tool acquisition. A team with one dashboard and a strict P99 SLO will outperform a team with five tools and no threshold, every time.

— Usama

Get engineering help that treats optimization as a discipline, not a patch

Bitrupt

If your team has read this far and recognized a gap between what your monitoring shows and what your users are actually feeling, that gap is exactly what a senior engineering team closes faster than a generalist one. The company staffs every engagement with senior engineers, working through flexible models like development pods or direct staff augmentation, to help performance audits progress efficiently. Clients get direct access to the engineers doing the work, with response times inside 24 hours, whether the job is a targeted optimization sprint or a full mobile app build from scratch.

For teams facing chronic crash clusters, unpredictable ANR spikes, or a cold startup time that affects retention, this is the kind of engagement this company runs regularly across several sectors. If your roadmap also includes AI-assisted monitoring or feature work, Bitrupt’s AI & Data team can scope that alongside a performance engagement rather than as a separate project. Start with a conversation about where your app’s P99 actually sits today.

Sources

FAQ

What does it mean to optimize an app?

Optimizing an app means systematically improving its speed, stability, and resource efficiency, covering startup time, frame rendering, memory use, crash rates, and battery consumption. It’s an ongoing engineering practice built on measurement, not a one-time cleanup task, and it directly affects retention and store visibility.

How do I turn off app optimization on Android?

Android’s system-level battery optimization settings can be adjusted per app under device battery settings, which controls how aggressively the OS restricts background activity for that app. This is separate from the engineering-side app optimization covered in this guide, which focuses on how the app itself performs rather than system battery restrictions.

How do I optimize my app?

Start by fixing your top crash and ANR clusters, since Android vitals ties both directly to Play Store visibility. From there, generate a Baseline Profile to cut cold startup latency, and set up continuous monitoring with tools like Macrobenchmark, Perfetto, and Firebase Performance Monitoring so regressions get caught before users feel them.

What is the best app to optimize my phone?

Phone-level cleaner and optimizer apps generally offer limited, temporary improvements and aren’t a substitute for the engineering-side optimization work covered here, which happens inside the app’s own codebase. If you’re a developer trying to improve your own app’s performance, the tools that matter are platform-native ones like Android Studio Profiler, Macrobenchmark, and Xcode Instruments, not a third-party phone cleaner.

End of essay
Rate this essay

Was this
worth your time?

One tap. No signup, no mailing list — just a signal that helps us write the next one better.

Tap a star
Start a project
Tell us what you’re building.We’ll ship it.

Send a few details and a senior engineer — not a sales rep — gets back to you with a clear next step within a day. In a hurry? .

+1 (302) 899-1332Call us direct · US line
NDA-friendlyYour idea and IP stay 100% yours.
Reply within 24hA senior engineer, not a sales bot.
United States · Registered office8 The Green, Suite B, Dover, DE 19901+1 (302) 899-1332
PakistanOffice No 115, First Floor, SIDCO Avenue Center, Saddar, Karachi+92 312 282-8442
Prefer email?contact@bitrupt.co
+1

By submitting you agree to our privacy policy. We’ll never share your details.