Cloudflare Fixes Flaw Leaking Data Across Customer Containers

Cloudflare Fixes Flaw Leaking Data Across Customer Containers

Enterprise security teams are now evaluating the implications of this storage flaw on compliance-sensitive workloads hosted within serverless container environments. The discovery of a vulnerability in a major edge computing platform underscores the persistent difficulty of maintaining perfect isolation in multi-tenant architectures. While cloud providers often tout the robustness of their sandboxing technologies, this specific incident revealed that the underlying storage management systems can sometimes leave digital footprints that cross tenant boundaries. The vulnerability did not rely on sophisticated bypasses of modern encryption or complex social engineering; instead, it exploited a fundamental oversight in how shared disk space was recycled and provisioned to subsequent users.

This discovery has sparked a broader conversation regarding the shared responsibility model and the degree of transparency required when infrastructure-level flaws occur. Cloudflare disclosed that the gap in its Containers platform permitted one paying customer to read data remnants left behind by a completely different customer’s earlier workload. The flaw was brought to light in late September 2026, roughly three weeks after a security researcher identified the issue through a bug bounty program. It serves as a reminder that as infrastructure becomes more ephemeral and shared, the mechanisms for cleaning up after a process finishes must be just as high-speed and reliable as the deployment itself.

1. Public Announcement: Cloudflare’s Disclosure

The formal disclosure of the vulnerability on September 24, 2026, sent ripples through the cybersecurity community, particularly among those who rely on serverless containers for processing high-value data. Cloudflare confirmed that the flaw resided within its Containers platform, which was designed to allow Workers customers to execute full containerized workloads at the edge. The nature of the bug was particularly concerning because it required no authentication bypass or stolen cryptographic keys. Instead, a simple, aligned write operation to a shared disk was sufficient to retrieve several dozen kilobytes of a stranger’s data, effectively turning a standard storage operation into an unintended information harvesting tool.

Security analysts noted that the disclosure arrived several weeks after the initial report, a delay often necessary to ensure a complete global cleanup before publicizing the vulnerability. The researcher who flagged the issue demonstrated that by writing a mere 4 KB of data to a specific spot on the disk, they could consistently trigger the system to return 60 KB of residual data from previous tenants. This low barrier to entry for exploitation emphasized the severity of the flaw, as it bypassed traditional security layers like identity and access management. The company took the step of publishing a detailed technical post-mortem to explain how such a fundamental storage oversight could persist in a production environment.

2. Sequence of Events: Detection to Resolution

The timeline of this incident began on September 4, 2026, when Oren Yomtov, a researcher at the firm Accomplish, submitted a detailed report through Cloudflare’s bug bounty program. Recognizing the gravity of a cross-tenant data leak, the engineering team prioritized a mitigation strategy immediately. Within three days, by September 7, the company had implemented a primary fix designed to block the specific exploit path identified by the researcher. This rapid initial response was critical to narrowing the window of exposure while a more comprehensive, fleet-wide remediation plan was formulated and executed across thousands of global edge nodes.

Following the initial mitigation, the engineering teams spent the subsequent week validating the fix and ensuring that no alternative methods could bypass the new protections. On September 14, Cloudflare confirmed that the original proof-of-concept provided by the researcher was no longer functional. However, the most labor-intensive phase began shortly thereafter, as the company had to rebuild and replace cached images and storage disks across its entire global network to purge any remaining data remnants. This massive infrastructure overhaul concluded on September 19, leading to the public disclosure five days later once the platform was deemed fully sanitized and secure for all users.

3. Standard Multi-Tenant Storage Isolation: The Role of DM-Thin

Modern cloud infrastructure relies on highly efficient storage allocation to handle thousands of concurrent customer workloads on shared physical hardware. To achieve this, many platforms utilize Linux’s device mapper thin provisioning system, commonly referred to as dm-thin. This technology allows providers to oversubscribe storage by only allocating physical blocks when a container actually writes data, rather than pre-allocating large, empty chunks of disk for every tenant. When a container is decommissioned, these blocks are returned to a shared pool, where they wait to be reassigned to the next requesting workload, a process that happens millions of times daily.

For this multi-tenant model to remain secure, the platform must guarantee that any block handed to a new user is completely “zeroed” or erased before they can interact with it. If the erasure step is skipped or incomplete, the new tenant might inherit the digital leftovers of the previous occupant. In high-performance environments, there is often a tension between the speed of provisioning a new container and the time required to perform a thorough wipe of the underlying storage. Security best practices dictate that no block should ever be reused without a verified clearing process, as even fragmented data can contain highly sensitive information like session tokens or configuration files.

4. Technical Root: The Skip Block Zeroing Oversight

The vulnerability at the heart of the Cloudflare Containers leak was traced back to a specific configuration setting within the dm-thin system known as skip_block_zeroing. This option is typically used in environments where performance is the primary concern, as it allows the system to bypass the time-consuming process of filling a storage block with zeros before handing it to a new user. By enabling this setting, Cloudflare intended to accelerate the startup times for its serverless containers, but in doing so, it inadvertently created a mechanism for data to persist across tenant boundaries. The system was configured to allocate storage in 64 KB units, which became the baseline for the potential data exposure.

The exploit mechanics were deceptively simple: when a new container workload performed a 4 KB write that was correctly aligned with the start of a block, the dm-thin system would allocate the required 64 KB physical block to that container. Because the zeroing process was skipped, only the 4 KB actually written by the new user was overwritten with their own data. The remaining 60 KB of that physical block remained untouched, containing whatever data had been written there by the last customer who used that specific part of the hardware. This architecture allowed a user to effectively peer into the past state of the shared disk through a legitimate write operation.

5. Impact of Small Writes: Massive Data Exposure

The disproportionate nature of the leak—where a tiny 4 KB write could reveal 60 KB of old data—made the flaw particularly potent. For an attacker, the efficiency of this method meant that they could potentially harvest large amounts of information by repeatedly triggering new block allocations. Cloudflare’s internal audit of the vulnerability confirmed that the residual data was not merely random digital noise. Instead, researchers and engineers discovered that the recovered blocks often contained structured data that could be easily interpreted or reconstructed by an adversary with basic forensic tools.

Among the findings during the investigation were intact directory structures, which could reveal the file organization of a previous user’s application. More critically, the audit found database pages and, in some instances, completely intact SQLite database files. The presence of such structured information suggests that the risk was not limited to transient logs but extended to the core operational data of the affected workloads. For businesses running databases or complex file systems within these containers, the possibility that their entire data structure could be lifted by a subsequent user represented a catastrophic failure of the expected isolation boundaries.

6. Shared Infrastructure: Vulnerability in Cloudflare Sandboxes

The scope of the storage flaw was not confined solely to the Containers product; it also extended to Cloudflare Sandboxes. This secondary service, built on the same underlying container infrastructure, was designed to execute untrusted code or AI agents in an isolated environment. Because Sandboxes and Containers shared the same pool of physical disks and the same storage management configurations, they both inherited the skip_block_zeroing flaw. This meant that a vulnerability discovered in one product line effectively compromised the data integrity of another, illustrating how deeply rooted infrastructure issues can propagate across a service provider’s entire ecosystem.

This cross-product impact highlighted a common challenge in modern cloud architecture: the use of a unified “base layer” for multiple specialized services. While this approach allows for rapid scaling and consistent performance, it also creates a single point of failure for security. Any customer utilizing Sandboxes for sensitive tasks, such as code interpretation or temporary data processing, faced the same risk of their data leaking to a Containers user, and vice versa. The realization that the vulnerability was platform-wide forced Cloudflare to conduct a much broader remediation effort than would have been required for a standalone software bug.

7. Findings of the Internal Audit: A Global Footprint

Once the mechanism of the leak was understood, Cloudflare launched a comprehensive internal audit to determine the actual prevalence of residual data across its network. The results were startling: out of 24 tested placements, 18 showed evidence of residual data belonging to other tenants. This indicates a high probability that any customer using the service during the period of vulnerability had their data remnants sitting on a shared disk, waiting to be potentially accessed by the next occupant. The audit demonstrated that the flaw was not an isolated incident or a rare edge case, but a systemic condition of the platform’s storage logic.

The geographic reach of the vulnerability was equally extensive, with the audit identifying affected nodes across four different continents. This confirmed that the configuration error was part of the standard deployment template for Cloudflare’s global fleet, rather than a localized misconfiguration. Furthermore, the investigation examined over 5,600 directory blocks and found thousands of foreign directory markers. This data provided a concrete map of how widespread the cross-tenant exposure had become, proving that the risk was active in nearly every corner of the company’s edge network where these specific container services were operational.

8. Exploitation Difficulty: Hurdles for Potential Attackers

Despite the severity of the flaw, there were significant practical hurdles that would have faced any malicious actor attempting to exploit it for a targeted attack. To trigger the leak, a user needed a paid Cloudflare account and the ability to deploy a container capable of performing raw disk operations. Furthermore, the attacker would have had to understand the specific block alignment requirements to ensure the dm-thin system allocated a new block without zeroing it. These requirements raised the barrier to entry above that of a typical web-based vulnerability, requiring a level of technical sophistication and a financial relationship with the provider.

The most significant defense against a targeted breach was the inherent randomness of Cloudflare’s workload placement. An attacker could not choose which physical host their container would land on, nor could they predict which customer had used that specific block of storage previously. This made it virtually impossible to target a specific company or individual. While an attacker could harvest random data from the “next available” block, they had no control over the quality or ownership of that data. This opportunistic nature of the bug meant that while it was a significant privacy and security failure, it was an inefficient tool for corporate espionage.

9. Remediation and Global Cleanup: Restoring the Perimeter

The process of fixing the vulnerability required a dual-track approach to ensure both immediate protection and long-term sanitization. The first and most critical step was a configuration change to disable the skip_block_zeroing setting across the entire fleet. By ensuring that every new block allocation was accompanied by a mandatory zeroing process, Cloudflare effectively closed the door on any new data leaks. This change was implemented rapidly following the researcher’s report, serving as a primary defense that protected all new workloads from inheriting legacy data.

However, simply changing the configuration did not address the data that was already sitting on disks in an un-zeroed state. To fully resolve the issue, Cloudflare had to undertake a massive infrastructure overhaul that involved retiring every container disk that had functioned under the old, insecure settings. This required the company to systematically cycle through its global edge nodes, deleting old disk images and rebuilding cached snapshots from scratch. Because this remediation took place at the infrastructure level, customers were not required to take any manual action, though the scale of the operation meant it took nearly two weeks to ensure every trace of residual data was purged.

10. Uncertainty Regarding Prior Exploits: The Logging Gap

Cloudflare has maintained that there is no evidence suggesting the vulnerability was exploited by anyone other than the reporting researcher. While this is a reassuring statement, it comes with a significant technical caveat regarding the limitations of standard cloud logging. Most platform-level logs are designed to track high-level events like API calls, authentication attempts, and network traffic. They are rarely configured to monitor the specific, low-level raw disk read and write patterns that would characterize an exploitation of a storage-zeroing flaw. Consequently, the absence of an alert in the logs does not definitively prove that no unauthorized access occurred.

For organizations with high security requirements, this “logging gap” remains a point of concern. The company’s audit confirmed that the data exposure was widespread, yet the tools to detect whether that exposure was leveraged by a malicious actor were not fully in place at the time. This uncertainty highlights a common issue in cloud security where the ability to detect a breach lags behind the existence of the vulnerability itself. Without granular telemetry at the storage layer, the true history of the flaw’s lifecycle remains partially obscured, leaving security teams to rely on the provider’s internal assessments and their own risk calculations.

11. Comparison to Other Cloud Security Incidents: Room Cleaning vs. Escape

The Cloudflare Containers flaw belongs to a different class of vulnerabilities than the “container escape” bugs that have dominated headlines in recent years. In a typical escape scenario, such as the famous Azurescape or GKE Fragnesia incidents, a malicious user breaks through the software boundaries of their container to gain access to the underlying host or other customers’ environments. These are active attacks on the isolation perimeter. In contrast, the Cloudflare incident was a passive data-remnant bug. The container remained perfectly confined within its assigned boundaries, but the “room” it was given by the host had not been cleaned of the previous occupant’s belongings.

This distinction is important for threat modeling. While escape bugs are often more technically complex and offer greater control to an attacker, storage-reuse flaws like this one are often easier to trigger and can be just as damaging in terms of data loss. The Cloudflare incident demonstrates that even if a container runtime is perfectly secure, the surrounding infrastructure—specifically the storage and memory management systems—must be equally scrutinized. It is a reminder that isolation is a multi-layered concept, and a failure in any single layer, even one as mundane as disk provisioning, can lead to a significant breach of tenant privacy.

12. Competitive and Market Consequences: The Trust Factor

Cloudflare’s decision to provide a detailed and transparent account of the vulnerability has been praised by some as a model for industry communication. However, the incident still presents a challenge to the company’s competitive standing against established giants like AWS, Google Cloud, and Microsoft Azure. These competitors have spent years refining their multi-tenant isolation technologies and have survived their own high-profile security scares. For Cloudflare, which is a relatively newer entrant into the serverless container market, a flaw of this nature could cause potential customers to re-evaluate the maturity of its infrastructure compared to more seasoned providers.

The market impact is likely to be felt most acutely in sectors with stringent compliance and privacy requirements, such as finance, healthcare, and government services. For these users, the mere possibility of cross-tenant data exposure can be a deal-breaker, regardless of how quickly the fix was deployed. While Cloudflare’s performance and edge-latency advantages remain strong, security teams in these industries often prioritize proven isolation track records over speed. This incident may prompt a more rigorous vetting process for all “edge container” services, as organizations weigh the benefits of decentralized computing against the risks of shared infrastructure vulnerabilities.

13. Guidance for Security and Compliance Departments: Actionable Steps

In the wake of this disclosure, security and compliance departments should take proactive steps to assess their exposure and harden their own practices. The first priority should be a retrospective audit of any sensitive data that was processed within Cloudflare Containers or Sandboxes during the risk window, which closed on September 19, 2026. While Cloudflare suggests no direct action is needed, organizations should verify whether any long-lived credentials, API keys, or proprietary data models were stored on local container disks. If such data was present, rotating those credentials and monitoring for unusual activity is a prudent defensive measure.

Beyond the immediate incident, this flaw should serve as a catalyst for updating vendor risk assessment protocols. Organizations should specifically ask their cloud providers about their storage allocation policies and how they guarantee the zeroing of blocks between tenants. Moving forward, it is a best practice to avoid storing permanent secrets or sensitive databases on the local, ephemeral disks of a container. Instead, teams should utilize external secret managers and encrypted object storage for any data that must persist or remains highly sensitive. By treating the local disk as an untrusted and public-facing resource, developers can minimize the impact of any future storage-layer vulnerabilities.

14. Future Outlook for Container Security: Evolving Industry Standards

The disclosure of the Cloudflare storage flaw signaled a shift in how both providers and regulators view the security of shared infrastructure. In the months following the incident, industry experts noted an increase in audits targeting similar storage configurations across the cloud ecosystem. It is highly probable that other providers will quietly revise their thin-provisioning settings to ensure that performance optimizations never come at the expense of tenant isolation. This event has also likely increased the “market value” of isolation research, leading to higher bug bounty payouts for researchers who can identify similar flaws in memory or storage management.

From a regulatory perspective, there is growing interest in the systemic risks posed by the shared nature of cloud infrastructure. Future standards may mandate that cloud providers offer more granular proof of data destruction and isolation, perhaps through audited, automated reports that verify block-zeroing at the hardware level. As the technology matures, the industry will likely move toward a model where isolation is not just a promise, but a verifiable state. The lessons learned from this vulnerability ensured that the next generation of serverless container platforms would be built with a deeper focus on the fundamental “cleanliness” of shared resources, ultimately leading to a more resilient and secure digital landscape for all enterprises.

Subscribe to our weekly news digest.

Join now and become a part of our fast-growing community.

Invalid Email Address
Thanks for Subscribing!
We'll be sending you our best soon!
Something went wrong, please try again later