AI Sandboxes for Compliance-Bound Industries

There is a particular kind of vertigo that sets in when you're sitting across from a compliance officer who has just been told her organization's AI model is already in production, handling patient triage decisions, and nobody documented anything. The model works. The outcomes look fine. And yet the exposure is total, because the work of demonstrating that it works was never done. That scenario is no longer hypothetical in sectors where AI adoption has outrun governance by a margin that should concern anyone responsible for either. A 2025 HFMA report found that 88% of health systems are already using AI in some form, while only 18% have mature governance structures. That 70-point spread is not a governance gap; it is a liability corridor.
Regulatory sandboxes are the structural response to that corridor, not a workaround around it, and not a substitute for regulation. The concept is roughly a decade old, borrowed from fintech: the UK launched the first regulatory sandbox in 2016, initially to allow financial technology firms to test products under supervised rule relaxation before permanent frameworks applied. The logic translated naturally to AI because the underlying problem is identical. Technology development moves faster than formal rulemaking can keep pace with, and in compliance-bound industries, the cost of deploying without structured validation is not a bad user experience. It is legal exposure, patient safety risk, and reputational damage that compounds across regulatory cycles.
Understanding what sandboxes actually are, how they function, and what they require organizationally is the practical foundation before any regulated firm decides whether and how to enter one.
The Two Sandbox Archetypes and What Separates Them
The term "sandbox" is used loosely enough in practice that conflating its two principal forms creates real strategic confusion. They are different mechanisms, serving different purposes, and regulated organizations typically need both, sequenced correctly.
The first archetype is the regulatory sandbox: a government- or regulator-supervised program that temporarily relaxes or waives specific rules so firms can test AI systems before permanent regulatory frameworks fully apply. Participation confers formal legal status. The firm is operating under regulatory guidance, not in spite of it. The World Economic Forum's 2025 reporting on AI governance categorizes these supervised, government-collaborative environments as a distinct type within a broader taxonomy, precisely because the defining feature is the relationship with the regulator, not the technical infrastructure. Entry is competitive. Oversight is ongoing. Exit produces documentation.
The second archetype is the enterprise sandbox: a firm-operated internal environment, typically using synthetic or anonymized data, where models are tested before internal deployment. No external regulator is involved by default. The firm sets the rules, the guardrails, and the exit criteria. This is where most regulated organizations begin, even those that eventually enter a formal regulatory program. It is also where most stop, which is a strategic mistake for organizations facing high-risk AI deployment decisions.
A third dimension cuts across both archetypes: infrastructure. On-premises sandboxes offer full sovereignty over hardware, data, and security, which matters critically when data cannot legally leave a controlled environment. Cloud sandboxes offer agility and scalability, but require robust data governance to satisfy regulators. Most enterprises now arrive at a hybrid approach, using cloud infrastructure for development speed and on-premises environments for the most sensitive compliance workloads. The infrastructure choice is not merely a technical decision; it is a compliance decision, and it should be treated as one from the beginning.
The distinction that matters operationally: regulatory sandboxes are about legal status and dialogue with regulators. Enterprise sandboxes are about technical readiness. One does not substitute for the other.
How a Regulatory Sandbox Actually Operates: Entry, Oversight, and Exit
The mechanics of a regulatory sandbox are more structured than most organizations expect when they first encounter the concept. Entry is not a formality.
Applicants must identify the AI system and its intended use, specify the regulatory rules that would ordinarily apply, and articulate why guidance or temporary relaxation is being sought. Most programs require applicants to name specific risks to consumers, health, or safety, and to submit mitigation plans alongside that identification. This is not optional documentation; it is the application. Spain's first sandbox cohort admitted 12 projects across six sectors, meaning selection is competitive and scope is defined. Firms that arrive without having done the risk-identification homework typically do not get in.
Once inside, the sandbox operates as a supervised process, not a one-time approval. The firm receives ongoing regulatory dialogue as it tests, with guidance that can adapt as the model's behavior is observed. Synthetic or anonymized data is used to satisfy privacy requirements while still generating test results that carry analytical weight. Documentation runs throughout the entire engagement, and that documentation is not merely a record for internal purposes. Under the EU AI Act framework, it becomes the compliance record firms use to demonstrate conformity later. It is admissible evidence, not a rehearsal.
The liability structure deserves particular attention because it is frequently misunderstood. Firms that follow the guidance of the national competent authority during sandbox participation are protected from administrative fines for infringements that occur during testing. That protection, however, does not extend to harm caused to third parties. If an AI system operating in a sandbox causes damages to individuals, the firm retains legal exposure. Regulatory protection and freedom from liability are not the same thing, and treating them as equivalent is a risk the legal and compliance functions of any regulated organization should be careful to avoid.
Exit is time-limited by design. Utah's program grants up to two years of what its legislation calls "regulatory mitigation"; EU programs similarly define an end date. Exit produces a compliance documentation record, formal regulator feedback, and in some cases broader published guidance. Spain's regulatory body, AESIA, published 16 practical compliance guides in December 2025 derived directly from sandbox experience, turning supervised experimentation into reusable market infrastructure. Exit is not automatic approval for deployment, but for high-risk AI systems, it substantially compresses the path to it.
Where Regulatory Sandboxes Exist Today, and Where the Gaps Are
The Datasphere Initiative identified more than 60 sandboxes worldwide related to AI, data, and technology as of early 2025, of which 31 are national sandboxes focused specifically on AI innovation including machine learning. The phenomenon is real and growing. It is also deeply uneven in ways that matter for where regulated organizations can actually access formal supervision.
Active programs operate in Spain, the UK, Singapore, Malaysia through Bank Negara, and in Utah in the United States. Southeast Asia has moved faster than most regions on financial-services AI sandboxes, with Singapore, Malaysia, and Thailand each developing distinct programs. Brazil and Kenya have notable programs, which is important context for understanding that this model is not exclusive to wealthy jurisdictions. China has promoted sandboxes for AI best practice, with documented applications in SME financing and anti-money laundering.
The EU implementation picture is the most consequential example of structural lag. As of August 2025, only one of the EU's 27 member states had an operational AI regulatory sandbox; five were actively implementing; four had declared intent; and 16 had not communicated plans. In May 2026, the EU extended the deadline for member states to establish sandboxes from August 2026 to August 2027, a one-year extension granted at a moment when the gap between mandate and operational reality was already stark. A proposed Digital Omnibus amendment would create an EU-level sandbox operating alongside national ones, a dual-layer structure that could allow firms to engage with the EU AI Office directly rather than only through national regulators.
The United States presents a different structural challenge: state-led experimentation without a federal floor. Utah was first with an AI-specific sandbox under its 2024 AI Policy Act. Texas passed but did not enact its Responsible AI Governance Act in 2025, the bill was vetoed by the governor. Arizona, Wyoming, Florida, and North Carolina operate related programs. Federal proposals exist, including the SANDBOX Act and a July 2025 AI Action Plan recommending agency-level sandboxes, but no federal program is operational. The practical implication is that which sandbox is available to any given regulated organization depends heavily on where it is chartered and which sector it operates in.
Spain and Singapore as the Most Instructive Live Cases
Among all active programs, two stand out not simply for their maturity but for what they reveal about how sandboxes function in practice and what they produce.
Spain
Spain launched its AI regulatory sandbox under Royal Decree 817/2023, becoming operative in November 2023 and remaining the only EU member state with an operational sandbox as of August 2025. The first cohort admitted 12 projects spanning healthcare diagnostics, financial-services risk assessment, employment-related AI, biometrics, critical infrastructure, and machinery. That cross-sector composition was deliberate; the goal was to generate transferable learning, not sector-specific precedent alone.
By 2026, the program was in its third cohort, having processed more than 20 AI systems end-to-end. The most concrete output for other regulated firms is the 16 practical compliance guides AESIA published in December 2025, drawn directly from sandbox experience. Those guides represent something worth examining carefully: a sandbox run with sufficient rigor generates compliance infrastructure for the entire market, not only for the firms that participated. Firms that could not access the first cohort benefited from the institutional learning of those who did. That is a model outcome, and it points to a question regulated organizations should ask of any sandbox they consider: does participation contribute to shared compliance infrastructure, or does it only benefit the participant?
Singapore
Singapore's approach is instructive precisely because it does not rely on any single mechanism. The Monetary Authority of Singapore published AI risk management guidelines for financial institutions in November 2025, covering the full model lifecycle from development through retirement. The Veritas framework involved seven major financial institutions, including BNY Mellon, DBS, HSBC, OCBC, Singlife, Standard Chartered, and UOB, piloting the integration of FEAT principles with existing governance frameworks. A Global AI Assurance Sandbox launched in July 2025, designed specifically for generative AI systems.
DBS Bank's trajectory through this governance ecosystem is the most data-rich illustration of what sustained, governance-backed AI deployment can produce: $750 million in reported AI value in 2024, a projection exceeding $1 billion for 2025, and more than 350 use cases spanning fraud detection, credit assessment, and customer service. Those numbers are not a product of regulatory luck. They reflect years of governance investment that preceded and enabled scale.
Both cases point to the same observation: the sectors most represented in active sandboxes, healthcare, financial services, and insurance, are precisely the ones facing the largest governance gaps in the aggregate data. That convergence is not accidental.
What the EU AI Act Requires of Sandbox Participants, and Why It Matters Beyond Europe
Article 57 of the EU AI Act sets out the purposes sandboxes are meant to serve: improving legal certainty around compliance, sharing best practices, fostering innovation, contributing to evidence-based regulatory learning, and accelerating market access, particularly for SMEs and startups. The rationale is deliberately multi-stakeholder, which is why the compliance obligations attached to participation are worth understanding precisely rather than approximately.
Testing must be supervised by the national competent authority. Self-certification is not sufficient. High-risk AI systems are the primary focus, meaning the Act's tiered risk classification determines which systems face the most scrutiny and which are eligible for sandbox testing under its framework. Documentation must run throughout the engagement, and that documentation constitutes the compliance record.
The liability structure, as noted earlier, is precise. Administrative fine protection applies when firms follow competent authority guidance during experimentation. Third-party liability for harm caused during testing is not extinguished by sandbox participation. Legal and risk teams evaluating participation on behalf of regulated organizations need to hold both of those points simultaneously rather than reading the protection broadly.
For non-European organizations, the extraterritorial reach of the EU AI Act changes the calculus significantly. The Act applies to any AI system placed on the EU market, regardless of where the developing organization is based. A healthcare AI firm operating primarily in the United States but serving European patients or hospital systems faces EU AI Act requirements on those systems. Building compliance practices aligned with the Act's sandbox framework is not Europe-specific preparation; it is an investment with broader regulatory utility. The Act's framework is the most detailed in force globally, and other jurisdictions are developing against it or in active dialogue with it.
What Regulated Organizations Need in Place Before Entering or Building a Sandbox
The organizations that exit sandbox programs with the most to show for it are not necessarily those with the most sophisticated AI. They are the ones that arrived with organizational infrastructure sufficient to make the engagement productive. That infrastructure does not emerge during sandbox participation; it has to exist before it begins.
The Internal Governance Baseline
A defined model inventory is the starting point: knowing which AI systems are in development or deployment, how they are classified by risk, and who owns each one at the enterprise level. This is not a spreadsheet exercise. It is a structural requirement, because the sandbox application process requires firms to specify which system is being tested, what rules apply to it, and what the genuine regulatory uncertainties are. Organizations that lack a model inventory cannot complete that specificity.
Data governance is a technical prerequisite, not merely a policy position. The ability to produce synthetic or anonymized datasets that satisfy privacy regulations without sacrificing the test validity needed to generate meaningful results requires engineering investment. It cannot be improvised during sandbox participation.
Clear ownership means someone accountable for AI risk at the enterprise level, not only at the project level. Sandbox programs, particularly regulatory ones, require a single point of institutional accountability. Organizations that distribute AI risk ownership diffusely across project teams tend to struggle with the documentation continuity that both entry and exit require.
Infrastructure Considerations
The on-premises versus cloud versus hybrid question is, as mentioned, a compliance decision dressed in technical language. For regulated data categories, patient records and financial account data being the clearest examples, on-premises sovereignty is often non-negotiable at the testing stage. Regulators want to know that the data used in testing cannot leave a controlled environment without appropriate governance. Hybrid architectures, using cloud for development speed and on-premises for the most sensitive workloads, represent the practical resolution most enterprises arrive at after attempting one or the other exclusively.
Sequencing
The sequencing mistake organizations make most often is attempting to enter a regulatory sandbox before internal testing has validated the model's basic behavior. The enterprise sandbox should come first, precisely because it is the stage where the model's fundamental characteristics are established and documented. Regulatory observers should not be the first audience encountering a model that has not yet been tested internally.
Governance documentation should begin at the internal stage, not at the point of regulatory engagement. The Spain and Singapore cases illustrate this clearly: the compliance record built during testing is what carries value at exit. Firms that begin documenting only when regulators are watching have already lost the most credible portion of the evidentiary timeline.
The broader observation, drawn from watching organizations approach this across sectors, is that those who treat sandbox participation as a compliance checkbox tend to exit without the reusable infrastructure that would make the effort worthwhile for the next system they develop. The organizations that get lasting value from sandbox engagement are those that treat it as a governance-building exercise from the beginning, one that produces documentation, institutional knowledge, and regulator relationships that compound across future deployments. That compounding effect is what Spain demonstrated at the market level with its published guides. Individual organizations can replicate the same logic internally. Whether they do depends almost entirely on what they decided the sandbox was for before they entered it.


