{"id":5394,"date":"2026-07-23T12:06:32","date_gmt":"2026-07-23T06:36:32","guid":{"rendered":"https:\/\/www.getpanto.ai\/blog\/?p=5394"},"modified":"2026-07-23T12:27:12","modified_gmt":"2026-07-23T06:57:12","slug":"flaky-test-statistics","status":"publish","type":"post","link":"https:\/\/www.getpanto.ai\/blog\/flaky-test-statistics","title":{"rendered":"Flaky Test Statistics 2026: Rates, Causes &#038; Cost Data"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Google has reported that <strong>16% of its 4.2 million tests show some level of flakiness<\/strong>, and that <strong>84% of pass-to-fail transitions in its CI system are flaky, not real bugs<\/strong>. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Across the industry, Bitrise&#8217;s 2025 Mobile Insights report found the share of teams experiencing <a href=\"https:\/\/www.getpanto.ai\/blog\/best-flaky-test-tools\">test flakiness grew <\/a>from 10% in 2022 to 26% in 2025.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Flaky tests are automated tests that pass and fail on the same code without any change. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">They are one of the most measurable, best-studied reliability problems in software engineering, with data going back to Google&#8217;s original 2016 research and continuing through peer-reviewed studies published this year.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This article compiles verified flaky test statistics for 2026, covering prevalence, root causes, financial and time cost, detection accuracy, AI-based repair, and market growth.<\/p>\n\n\n<h2 class=\"wp-block-heading\" id=\"flaky-test-statistics-key-insights-and-takeaways\"><span class=\"ez-toc-section\" id=\"flaky-test-statistics-key-insights-and-takeaways\"><\/span>Flaky Test Statistics: Key Insights and Takeaways<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n<ul class=\"wp-block-list\">\n<li><strong>16% of Google&#8217;s 4.2 million tests show some level of flakiness<\/strong>, based on Google&#8217;s own published CI research. <br><br><\/li>\n\n\n\n<li><strong>Teams experiencing test flakiness rose from 10% in 2022 to 26% in 2025<\/strong>, a 160% increase, according to Bitrise&#8217;s analysis of over 10 million builds.<br><br><\/li>\n\n\n\n<li><strong>Atlassian estimates 150,000 developer hours are lost per year<\/strong> to flaky tests across its engineering organization. <br><br><br><strong>Slack cut its test-related CI failure rate from 56.76% to 3.85%<\/strong> after investing in dedicated flaky test remediation. <br><br><br><strong>45% of flaky test fixes address async and timing issues<\/strong>, the single largest root cause identified in academic research. <br><br><\/li>\n\n\n\n<li><a href=\"https:\/\/www.getpanto.ai\/products\/self-healing-test-automation\"><strong>AI-based repair tools<\/strong><\/a><strong> now fix 47.6% of reproducible flaky tests automatically<\/strong>, with over half of those fixes accepted by developers.<br><br><\/li>\n<\/ul>\n\n\n<h3 class=\"wp-block-heading\" id=\"at-a-glance-flaky-test-statistics-2026\"><span class=\"ez-toc-section\" id=\"at-a-glance-flaky-test-statistics-2026\"><\/span>At A Glance: Flaky Test Statistics 2026<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Metric<\/th><th>Figure<\/th><\/tr><\/thead><tbody><tr><td>Google test flakiness rate<\/td><td>16% of all tests<\/td><\/tr><tr><td>Google pass-to-fail transitions that are flaky<\/td><td>84%<\/td><\/tr><tr><td>Microsoft build flakiness rate<\/td><td>26% of sampled builds<\/td><\/tr><tr><td>Teams experiencing flakiness, 2022 to 2025<\/td><td>10% to 26%<\/td><\/tr><tr><td>GitHub commits hitting a flaky red build (2020)<\/td><td>9% (1 in 11)<\/td><\/tr><tr><td>Atlassian developer hours lost annually<\/td><td>150,000 hours<\/td><\/tr><tr><td>Slack CI failure rate before and after remediation<\/td><td>56.76% to 3.85%<\/td><\/tr><tr><td>Top root cause of flaky test fixes<\/td><td>45% async\/timing issues<\/td><\/tr><tr><td>AI repair rate for reproducible flaky tests<\/td><td>47.6%<\/td><\/tr><tr><td>Meta E2E test flakiness vs. unit tests<\/td><td>~10% vs. under 1%<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n<h2 class=\"wp-block-heading\" id=\"flaky-test-statistics-a-deep-dive\"><span class=\"ez-toc-section\" id=\"flaky-test-statistics-a-deep-dive\"><\/span>Flaky Test Statistics: A Deep Dive<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n<h3 class=\"wp-block-heading\" id=\"1-flaky-test-statistics-prevalence-statistics\"><span class=\"ez-toc-section\" id=\"1-flaky-test-statistics-prevalence-statistics\"><\/span>1. Flaky Test Statistics: Prevalence Statistics<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\">Flakiness is not a niche problem confined to a few unlucky teams. It shows up consistently across companies of very different sizes and testing maturity levels.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Google&#8217;s foundational research, first published by John Micco in 2016 and still widely cited today, found that <strong>almost 16% of Google&#8217;s 4.2 million tests showed some level of flakiness<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Google also reported that developers on its platform spend <strong>between 2% and 16% of total compute resources<\/strong> simply re-running flaky tests to confirm results.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Microsoft&#8217;s internal research identified flaky failures in <strong>26% of sampled builds<\/strong> across its large-scale CI systems, a figure cited in multiple later academic papers, including the 2024 FlaKat study.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">GitHub&#8217;s own 2020 engineering data showed that <strong>1 in 11 commits (9%)<\/strong> triggered at least one red build caused specifically by a flaky test, not a real code issue.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Source<\/th><th>Flakiness Metric<\/th><th>Value<\/th><th>Year<\/th><\/tr><\/thead><tbody><tr><td>Google<\/td><td>Tests with some flakiness<\/td><td>16%<\/td><td>2016<\/td><\/tr><tr><td>Google<\/td><td>Pass-to-fail transitions that are flaky<\/td><td>84%<\/td><td>2016<\/td><\/tr><tr><td>Microsoft<\/td><td>Sampled builds with flaky failures<\/td><td>26%<\/td><td>Cited 2024<\/td><\/tr><tr><td>GitHub<\/td><td>Commits with a flaky red build<\/td><td>9%<\/td><td>2020<\/td><\/tr><tr><td>Open source study (1,960 Java projects)<\/td><td>Projects affected by flaky builds<\/td><td>51.28%<\/td><td>Feb 2025<\/td><\/tr><tr><td>Same study<\/td><td>Rerun builds showing flaky behavior<\/td><td>67.73%<\/td><td>Feb 2025<\/td><\/tr><tr><td>Same study<\/td><td>Total builds that are rerun<\/td><td>3.2%<\/td><td>Feb 2025<\/td><\/tr><tr><td>Slack<\/td><td>CI failures from test job failures (before fix)<\/td><td>56.76%<\/td><td>2022<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">A separate 2025 academic study analyzed <strong>1,960 open source Java projects using GitHub Actions<\/strong> and found that <strong>51.28% of all projects were affected by flaky builds<\/strong>, with <strong>67.73% of rerun builds<\/strong> showing flaky behavior once teams triggered a retry.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Industry-wide, the trend is moving in the wrong direction. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Bitrise&#8217;s 2025 Mobile Insights report, based on <strong>over 10 million CI builds tracked across 3.5 years<\/strong>, found that the proportion of teams experiencing test flakiness climbed from <strong>10% in 2022 to 26% in 2025<\/strong>,.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is a 160% increase, while <a href=\"https:\/\/www.getpanto.ai\/blog\/detect-flaky-tests\">CI pipeline complexity<\/a> grew 23% over the same period.<\/p>\n\n\n<h3 class=\"wp-block-heading\" id=\"2-flaky-test-statistics-causes-statistics\"><span class=\"ez-toc-section\" id=\"2-flaky-test-statistics-causes-statistics\"><\/span>2. Flaky Test Statistics: Causes Statistics<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\">Understanding what actually causes flaky tests matters more than knowing they exist. Academic research has isolated specific, repeatable categories rather than treating flakiness as random noise.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The most cited study on flaky test causes comes from Luo et al. (FSE 2014), which analyzed <strong>201 flaky test fixes across 51 Apache open source projects<\/strong>. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That research found <strong>45% of all fixes addressed asynchronous wait and timing issues<\/strong>, making it by far the single largest cause category.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A later ICSE 2021 study focused specifically on UI-driven tests and found that <strong>async-wait issues accounted for roughly 45% of UI-specific flaky failures<\/strong> as well, reinforcing timing as the dominant root cause across both backend and frontend test types.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Root Cause Category<\/th><th>Share of Flaky Test Fixes<\/th><\/tr><\/thead><tbody><tr><td>Async wait \/ timing issues<\/td><td>45%<\/td><\/tr><tr><td>Concurrency (data races, deadlocks)<\/td><td>16% of studied commits<\/td><\/tr><tr><td>Test order dependency<\/td><td>9% of studied commits<\/td><\/tr><tr><td>Resource leak<\/td><td>5% of studied commits<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Owain Parry&#8217;s survey of flaky test research (published in ACM TOSEM, 2021) categorized causes across multiple open source codebases, including the Home Assistant project. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That analysis found <strong>16% of relevant commits fell under the concurrency category<\/strong>, covering thread interaction issues like data races and deadlocks, with a further <strong>9% under test order dependency<\/strong> and <strong>5% under resource leaks<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A separate finding from Parry&#8217;s 2025 follow-up research adds an operational insight: <strong>roughly 75% of flaky tests cluster around a shared underlying root cause<\/strong>, meaning fixing one infrastructure or timing issue can resolve more than a dozen flaky tests at once rather than requiring test-by-test fixes.<\/p>\n\n\n<h3 class=\"wp-block-heading\" id=\"3-flaky-test-statistics-time-and-compute-statistics\"><span class=\"ez-toc-section\" id=\"3-flaky-test-statistics-time-and-compute-statistics\"><\/span>3. Flaky Test Statistics: Time and Compute Statistics<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\">Flaky tests carry a real, quantifiable cost in both engineering hours and compute spend, not just an abstract drag on morale.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Atlassian&#8217;s engineering team estimated in 2025 that flaky tests cost the organization <strong>150,000 developer hours per year<\/strong>. That figure comes from tracking investigation time, re-runs, and delayed merges across its internal CI systems.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A 2024 industrial case study, published at ICST 2024 by Leinen et al., measured actual developer time spent at a mid-sized engineering organization. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It found a team of roughly 30 developers spent <strong>2.5% of total productive time<\/strong> dealing with flaky tests, including <strong>1.3% specifically on repair work<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Slack&#8217;s engineering team published some of the clearest before-and-after data available. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Before investing in dedicated flaky test detection and suppression tooling, <strong>56.76% of Slack&#8217;s CI failures were test job failures<\/strong>, a combination of flaky and genuinely broken tests. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">After remediation, that number dropped to <strong>under 4% (3.85%)<\/strong>.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Company or Study<\/th><th>Cost Metric<\/th><th>Value<\/th><\/tr><\/thead><tbody><tr><td>Atlassian<\/td><td>Developer hours lost annually<\/td><td>150,000 hours<\/td><\/tr><tr><td>ICST 2024 case study<\/td><td>Productive time spent on flaky tests (30-dev team)<\/td><td>2.5%<\/td><\/tr><tr><td>Same study<\/td><td>Time spent specifically on repair<\/td><td>1.3%<\/td><\/tr><tr><td>Google<\/td><td>Compute resources spent re-running flaky tests<\/td><td>2% to 16%<\/td><\/tr><tr><td>Slack<\/td><td>CI failure rate before remediation<\/td><td>56.76%<\/td><\/tr><tr><td>Slack<\/td><td>CI failure rate after remediation<\/td><td>3.85%<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.getpanto.ai\/blog\/google-play-statistics\">Google&#8217;s own figures<\/a> on compute cost are notable at scale. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The company has reported spending <strong>between 2% and 16% of its total test compute resources<\/strong> purely on re-running tests to determine whether a failure was real or flaky.<\/p>\n\n\n<h3 class=\"wp-block-heading\" id=\"4-flaky-test-statistics-test-types\"><span class=\"ez-toc-section\" id=\"4-flaky-test-statistics-test-types\"><\/span>4. Flaky Test Statistics: Test Types<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\">Flakiness is not evenly distributed across test types. End-to-end and UI tests are consistently the least reliable category.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Meta&#8217;s engineering data shows that its <a href=\"https:\/\/www.getpanto.ai\/blog\/best-end-to-end-testing-tools\"><strong>end-to-end tests run at approximately 10% flakiness<\/strong><\/a>, while unit tests on the same codebases stay <strong>well under 1%<\/strong>. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That roughly tenfold gap reflects how much more surface area E2E tests expose to timing, network, and environment variance compared to isolated unit tests.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A Chromium CI study (Lampel et al., ESEC\/FSE 2023) offers a useful caution about detection accuracy at this test-type level. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Its flakiness prediction model reached <strong>99.2% precision<\/strong> in identifying flaky tests, yet still <strong>misclassified 76.2% of genuine fault-triggering failures as flaky<\/strong>. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The same study found an average of <strong>250 flaky tests per build<\/strong>, against just <strong>1 fault-revealing test per failing build<\/strong>, and that fault-revealing failures occurred in <strong>24.15% of all studied builds<\/strong>.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Test Type<\/th><th>Approximate Flakiness Rate<\/th><\/tr><\/thead><tbody><tr><td>Unit tests<\/td><td>Under 1%<\/td><\/tr><tr><td>End-to-end (E2E) tests<\/td><td>~10%<\/td><\/tr><tr><td>Chromium CI, average flaky tests per build<\/td><td>250<\/td><\/tr><tr><td>Chromium CI, fault-revealing tests per failing build<\/td><td>1<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n<h3 class=\"wp-block-heading\" id=\"5-flaky-test-statistics-ai-and-automated-flaky-test-detection-statistics\"><span class=\"ez-toc-section\" id=\"5-flaky-test-statistics-ai-and-automated-flaky-test-detection-statistics\"><\/span>5. Flaky Test Statistics: AI and Automated Flaky Test Detection Statistics<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.getpanto.ai\/blog\/best-flaky-test-tools\">AI-based tools are now being measured against flaky test detection<\/a> and repair with published, peer-reviewed benchmarks, not just vendor claims.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">FlakyGuard, presented at ASE 2025, demonstrated that AI could <strong>automatically repair 47.6% of reproducible flaky tests<\/strong>, with <strong>51.8% of those AI-generated fixes accepted by developers<\/strong> without further changes. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The approach treats test code as a graph structure and selectively explores relevant context rather than rewriting entire test files.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A separate ICSE 2024 study measured large language models specifically on categorized flaky test types. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It found LLMs successfully repaired <strong>79% of order-dependent flaky tests<\/strong> and <strong>58% of implementation-dependent flaky tests<\/strong>, showing meaningfully different success rates depending on the underlying cause.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.getpanto.ai\/blog\/llm-statistics\"><em>Check out our complete report on LLM Statistics \u2192<\/em><\/a><\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>AI Detection or Repair Study<\/th><th>Result<\/th><\/tr><\/thead><tbody><tr><td>FlakyGuard reproducible test repair rate<\/td><td>47.6%<\/td><\/tr><tr><td>FlakyGuard developer acceptance of AI fixes<\/td><td>51.8%<\/td><\/tr><tr><td>LLM repair rate, order-dependent flakes<\/td><td>79%<\/td><\/tr><tr><td>LLM repair rate, implementation-dependent flakes<\/td><td>58%<\/td><\/tr><tr><td>Chromium flakiness predictor precision<\/td><td>99.2%<\/td><\/tr><tr><td>Same model, real bugs misclassified as flaky<\/td><td>76.2%<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Google&#8217;s own root-cause localization research, tested across <strong>428 internal projects<\/strong>, reported <strong>82% accuracy<\/strong> in automatically identifying the code-level location responsible for a test&#8217;s flaky behavior, well before generative AI models entered the picture.<\/p>\n\n\n<h3 class=\"wp-block-heading\" id=\"6-flaky-test-statistics-market-trends\"><span class=\"ez-toc-section\" id=\"6-flaky-test-statistics-market-trends\"><\/span>6. Flaky Test Statistics: Market Trends<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\">The financial and tooling response to flaky tests is scaling alongside the problem itself. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Industry estimates place the <strong>AI-enabled software testing market at $1.01 billion in 2025, projected to reach $4.64 billion by 2034<\/strong>, a compound annual growth rate of roughly 18.3%. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This figure should be treated as a market research estimate rather than a company-reported number, and reflects the broader <a href=\"https:\/\/www.getpanto.ai\/products\/ai-automation-testing\">AI-assisted testing category<\/a> rather than flaky test tooling alone.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That growth tracks closely with the underlying problem. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Bitrise&#8217;s data shows CI pipeline complexity increased <strong>23% between 2022 and 2025<\/strong>, the same window in which team-level flakiness rates jumped from 10% to 26%. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As pipelines add more parallel jobs, more integrations, and more automated test suites, the raw surface area for flaky behavior expands with it.<\/p>\n\n\n<h3 class=\"wp-block-heading\" id=\"conclusion\"><span class=\"ez-toc-section\" id=\"conclusion\"><\/span>Conclusion<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\">The data is consistent across companies, academic studies, and years. Flaky tests affect somewhere between <strong>16% and 26% of tests or builds<\/strong> at major engineering organizations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">They cost real money too, <strong>150,000 lost developer hours<\/strong> at Atlassian alone. And they trace back overwhelmingly to <strong>async and timing issues<\/strong>, not exotic edge cases.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The trend line is also clear. Bitrise&#8217;s longitudinal data shows the share of teams affected grew <strong>160% between 2022 and 2025<\/strong>, even as CI tooling and AI-assisted testing matured.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For engineering and QA leaders, the numbers point toward the same conclusion. Flaky tests are not a gap that better discipline alone will close.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">They are a measurable, growing tax on CI infrastructure that scales with pipeline complexity. <a href=\"https:\/\/www.getpanto.ai\/\">AI-based repair and QA  tools<\/a> are starting to close part of that gap, fixing close to half of reproducible flaky tests automatically.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But detection accuracy studies like the Chromium research show these systems still misclassify a meaningful share of real bugs as noise. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The organizations narrowing the gap fastest, like Slack cutting failure rates from 56.76% to 3.85%, treat flakiness as a tracked, budgeted engineering metric rather than background noise.<\/p>\n\n\n<h3 class=\"wp-block-heading\" id=\"frequently-asked-questions\"><span class=\"ez-toc-section\" id=\"frequently-asked-questions\"><\/span>Frequently Asked Questions<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n<p class=\"wp-block-paragraph\"><strong>What percentage of tests are flaky?<\/strong> Google reports that 16% of its 4.2 million tests show some level of flakiness. Microsoft&#8217;s internal data found flaky failures in 26% of sampled builds, and an open source study of 1,960 GitHub Actions projects found 51.28% of projects affected by flaky builds.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>What is the biggest cause of flaky tests?<\/strong> Async and timing issues are the largest documented cause, responsible for 45% of flaky test fixes according to Luo et al.&#8217;s FSE 2014 study of 51 Apache projects, a figure echoed by a separate 2021 ICSE study on UI-specific flakiness.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How much do flaky tests cost engineering teams?<\/strong> Atlassian estimates 150,000 developer hours lost annually. A separate 2024 industrial case study found a 30-developer team spent 2.5% of total productive time on flaky tests, including 1.3% on repair work alone.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Can AI fix flaky tests automatically?<\/strong> Yes, with measured limits. FlakyGuard (ASE 2025) automatically repaired 47.6% of reproducible flaky tests, with 51.8% of fixes accepted by developers. LLM-based repair performs better on order-dependent flakes (79%) than implementation-dependent ones (58%).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Are flaky tests getting better or worse over time?<\/strong> Worse, according to the best available longitudinal data. Bitrise found the share of teams experiencing flakiness grew from 10% in 2022 to 26% in 2025, a 160% increase, even as CI tooling has improved.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Do all test types flake at the same rate?<\/strong> No. Meta&#8217;s data shows end-to-end tests flake at roughly 10%, compared to under 1% for unit tests, reflecting the larger surface area E2E tests expose to timing and environment variance.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Google has reported that 16% of its 4.2 million tests show some level of flakiness, and that 84% of pass-to-fail transitions in its CI system are flaky, not real bugs. Across the industry, Bitrise&#8217;s 2025 Mobile Insights report found the share of teams experiencing test flakiness grew from 10% in 2022 to 26% in 2025. [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":5396,"comment_status":"open","ping_status":"closed","sticky":false,"template":"wp-custom-template-panto-blogs-v3","format":"standard","meta":{"footnotes":""},"categories":[110],"tags":[],"class_list":["post-5394","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-qa-testing"],"_links":{"self":[{"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/posts\/5394","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/comments?post=5394"}],"version-history":[{"count":3,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/posts\/5394\/revisions"}],"predecessor-version":[{"id":5399,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/posts\/5394\/revisions\/5399"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/media\/5396"}],"wp:attachment":[{"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/media?parent=5394"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/categories?post=5394"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.getpanto.ai\/blog\/wp-json\/wp\/v2\/tags?post=5394"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}