Every practice owner who has ever considered an AI tool runs into the same wall. To find out whether the software actually works, someone has to test it. And to test it properly, you have to feed it patient data. That is the moment the room goes quiet, because feeding real patient records into an unproven tool is exactly the kind of thing that turns into a HIPAA headline.
There is a way around this that most practice owners have never heard of, even though the largest health agencies in the United States have relied on it for years. It is called Synthea, and it solves the problem by generating patient records that look and behave exactly like real ones but belong to people who do not exist.
Synthea is a free, open-source tool built by MITRE, the not-for-profit organisation that runs federally funded research centres for agencies including the CDC, CMS, and the FDA. It produces complete, realistic medical histories for entirely fictional patients: their conditions, their medications, their lab results, their visits, their allergies, all of it. Because none of these people are real, none of the data is protected health information. You can copy it, email it, upload it to a new vendor's tool, and hand it to a developer without a single privacy control, because there is nothing to protect.
For a practice that wants to evaluate AI without betting the business on an untested tool, this changes what is possible. This guide explains what Synthea is, why it matters, and how it fits into a safe path to adopting AI in a medical setting, all in plain language.
| Metric | 2026 Figure |
|---|---|
| Average cost of a US healthcare data breach (IBM, 2025) | $7.42 million |
| Healthcare's rank for breach cost among all industries | #1, for the 15th year running |
| US patient records exposed in breaches (2025) | ~275 million |
| Share of large healthcare breaches caused by hacking / IT incidents | ~80% |
| Diseases and conditions modelled by Synthea | 100+ |
| Cost to use Synthea | $0 (open source, free) |
Why This Matters: The Real Cost of Testing With Real Patient Data
Before explaining what Synthea does, it is worth being honest about the problem it solves, because the problem is expensive.
When a covered entity, and that includes almost every medical practice, allows protected health information to be used or disclosed in a way HIPAA does not permit, the law presumes a breach has occurred unless you can document that the risk of compromise was low. Handing a live patient database to a new AI vendor to "see if it works" is precisely the kind of disclosure that creates that exposure. If that vendor's environment is not secured, or there is no Business Associate Agreement in place, or the data leaks, the liability lands on your practice, not theirs.
The numbers explain why caution is warranted. According to IBM's 2025 Cost of a Data Breach report, the average healthcare breach in the United States cost $7.42 million, keeping healthcare the single most expensive industry for breaches for the fifteenth consecutive year. Across 2025, roughly 275 million patient records were exposed, and close to 80% of large healthcare breaches now trace back to hacking and IT incidents rather than lost laptops or paper files.
The point is not to frighten anyone out of adopting AI. The point is that the testing phase, the part where you are still deciding whether a tool is any good, is the worst possible moment to put real records at risk. You have not yet confirmed the tool is trustworthy, and you are already handing it your most sensitive asset. Synthea removes that trade-off entirely.
The Safe Way to Test AI: Fake Patients That Behave Like Real Ones
Synthea (the name is a blend of "synthetic" and "healthcare") is a piece of software that generates a fake but realistic population of patients. You tell it how many patients you want and, if you like, which state or town they should live in, and it produces a full set of medical records for each one.
Crucially, these are not random rows of nonsense. Synthea uses what MITRE calls a module framework, where each disease and care pathway is modelled on real clinical guidelines, CDC and NIH statistics, and academic research. So a synthetic 58-year-old patient with type 2 diabetes will have a believable history: the diagnosis at a plausible age, the medications a real clinician would prescribe, the follow-up visits, the lab values drifting over time, the related conditions that tend to travel with diabetes. Synthea models more than 100 diseases and conditions, covering the most common reasons people visit primary care and the leading causes of serious illness.
The records come out in the exact formats real health IT systems use:
- HL7 FHIR (the modern standard that new healthcare software speaks, available in R4 and older versions)
- C-CDA (the document format used for exchanging clinical summaries between systems)
- Plain CSV (simple spreadsheet files for analysis or bulk loading)
That last detail is what makes Synthea genuinely useful rather than a novelty. Because the output matches real-world standards, you can load a Synthea population straight into a test copy of an EHR, feed it to an AI tool, or hand it to a developer, and every system treats it as if it were real clinical data. The only difference is that no human being is behind any of it.
It is free, open source, and maintained by MITRE with contributions from a wide community. You can even download a pre-generated sample of over 1,000 synthetic patients from the Synthea project site without running anything yourself.
The One Idea That Makes Synthea Powerful: There Is Nothing to Protect
Everything valuable about Synthea flows from a single fact. Synthetic data is not protected health information.
HIPAA, GDPR, and every comparable privacy regime exist to protect information about identifiable real people. A synthetic patient generated by Synthea is not a real person, is not a de-identified real person, and cannot be re-identified back to anyone, because there is no one to re-identify. The record was invented from statistical models, not derived from a real chart.
This has practical consequences that are hard to overstate:
- You do not need a Business Associate Agreement to share Synthea data with a vendor, because no PHI changes hands.
- You do not trigger a breach notification obligation if synthetic data leaks, because nothing protected was disclosed.
- You can freely email it, store it in the cloud, post it in a shared folder, or hand it to an outside developer with no special controls.
- You can keep it around indefinitely for repeated testing, without the retention and disposal rules that govern real records.
Think of the difference like using a crash-test dummy instead of a person. Car makers do not test airbags on volunteers. They build a dummy that behaves like a human body under impact, precisely so they can run the dangerous test as many times as they want without anyone getting hurt. Synthea is the crash-test dummy for your patient data. You get to run every risky experiment, over and over, and no real patient is ever exposed.
Where Synthea Actually Gets Used: Real Adoption
Synthea is not a hobby project. It is infrastructure that serious institutions rely on.
The Centers for Disease Control and Prevention (CDC) adopted Synthea for its Childhood Obesity Data Initiative, working with MITRE to extend the tool so it could generate synthetic paediatric records that follow realistic childhood growth curves. That is a federal public health agency using invented patients to build and test data systems it could never safely build on real children's records.
MITRE itself, the organisation behind Synthea, operates federally funded research and development centres for the CDC, CMS (the agency that runs Medicare and Medicaid), the FDA, and others. Synthea grew out of that world, where the need to test national-scale health IT without exposing real citizens is a daily reality.
Beyond government, Synthea is widely used across three groups:
- Software developers and health IT vendors, who need realistic data to build and test their products before any real patient ever touches them.
- Researchers, who use synthetic populations to develop and share analysis methods openly, without the legal barriers that lock away real datasets.
- Educators and trainers, who teach students and staff on lifelike records without privacy risk.
For a practice owner, the takeaway is simple. When the CDC and the agencies that regulate you trust synthetic data enough to build real systems on it, you can trust it enough to test a vendor's AI tool with it.
How Synthea Fits a Safe Path to Adopting AI
This is where Synthea becomes directly relevant to your practice, even though you will almost certainly never run the tool yourself. Synthea belongs in a larger discipline: never let an unproven system touch real patients until it has proven itself somewhere safe first.
We have learned this lesson the expensive way across client engagements, and it is worth sharing plainly. In our work deploying AI systems, the single most important safety practice is that nothing goes near live data or real patients until it has been fully exercised in an environment that mirrors production but contains no real information. We have been called in to clean up after AI tools that were tested directly against live systems and wrote bad data into a real records database, or triggered real notifications during what was supposed to be a quiet test. Every one of those incidents would have been caught harmlessly if the tool had first been run against a realistic but fake dataset. Synthea is exactly that dataset for healthcare.
Here is what a safe evaluation looks like in practice, and where Synthea sits in it:
1. Generate a synthetic population that looks like your practice. A technical partner runs Synthea to create fake patients that resemble your real patient mix in age, common conditions, and volume. Now you have a realistic test set that carries zero privacy risk.
2. Load that data into a test copy of your systems. The synthetic records go into a staging version of your EHR or a sandbox environment. If you use an open platform like OpenEMR, this is straightforward, because Synthea's FHIR and C-CDA output loads cleanly into standards-based systems.
3. Point the AI tool at the fake data first. Whatever you are evaluating, an intake assistant, a scribe, a billing helper, a chatbot, it meets the synthetic patients before it ever meets a real one. You watch how it behaves across hundreds of believable cases, including the messy ones, and you find its failure modes while the stakes are zero.
4. Only then, and carefully, move to real data. Once the tool has proven itself against synthetic patients, you graduate it to a limited, closely monitored trial with real records and every proper safeguard in place: a signed Business Associate Agreement, access controls, and audit logging.
This staged approach is the backbone of every responsible AI deployment we run, and it is covered in more depth in our AI deployment checklist for medical practices. Synthea is what makes the safe first stage possible.
Why the Testing Stage Is Non-Negotiable With AI Tools Specifically
There is a reason this matters more for AI than for ordinary software, and it is worth understanding even at a non-technical level.
Traditional software is predictable. Give it the same input and it produces the same output every time. You can test it once, confirm it works, and trust it to keep behaving. AI tools built on large language models, the technology behind most of the AI products being sold to practices today, do not work that way. They are, by their nature, probabilistic. The same question can produce slightly different answers depending on the version of the model, the exact wording, or the surrounding context. A tool can give a perfectly formatted, confident answer that happens to be wrong.
This is not a flaw a vendor can simply patch away. It is a property of the technology. We cover this reality in detail in our guide to how doctors are actually using AI in practice, and the practical consequence is this: you cannot judge an AI tool from a polished demo. You have to watch it handle a wide spread of realistic cases, including the unusual and awkward ones, before you can trust it. That requires a large, varied, realistic dataset to test against. Real patient data would be ideal for realism, but using it at this stage is exactly the risk we are trying to avoid. Synthea gives you the realism without the risk. You can generate thousands of varied synthetic patients and see how the AI copes with the full range, all before a single real record is involved.
Synthea and Compliance: HIPAA, GDPR, and Beyond
Practice owners rightly want to know how a tool like this sits with the rules they live under. The good news is that Synthea's relationship with compliance is refreshingly simple, but there are two nuances worth understanding.
HIPAA. Synthetic data generated by Synthea is not PHI, so the HIPAA Privacy and Security Rules do not apply to it. You can use it, share it, and store it freely. The important nuance: this protection applies to Synthea's own generated data, not to any real data you might later mix in. The moment you move from synthetic testing to a real-patient trial, full HIPAA obligations return, and you need the usual safeguards. Synthea makes the early stage free of HIPAA burden; it does not remove HIPAA from the rest of your operation. Our guide on whether ChatGPT is HIPAA compliant explains where those obligations kick in for the AI tools themselves.
GDPR and international privacy law. For any practice or partner touching EU or UK patients, GDPR applies to personal data of identifiable living people. Truly synthetic data that cannot be linked to any real individual falls outside that definition, which is why synthetic data has become a recognised technique for privacy-preserving development in Europe as well. The same one caveat holds: the protection covers the invented data, not the real data you eventually test with.
A subtle point worth flagging. Not all "synthetic" data is equally safe. Some methods generate fake data by learning patterns from a real dataset, and if done carelessly, traces of real individuals can occasionally leak through. Synthea sidesteps this concern because it does not learn from a specific real patient database at all. It builds patients from published clinical models and population statistics, so there is no real chart sitting behind any synthetic record to leak from. For a compliance-conscious practice, that is a meaningful distinction, and it is worth asking any vendor exactly how their test data was produced.
What Synthea Is Not: Setting Honest Expectations
To use Synthea well, it helps to be clear about its limits.
It is not a consumer app. There is no friendly dashboard your office manager logs into. Synthea is a developer tool. In practice, your technical partner runs it and hands you the resulting data or loads it into your test systems. You benefit from it without ever operating it.
It is not a substitute for real-world validation. Synthetic patients are realistic, but they are generated from statistical models, so they will not perfectly capture every quirk of your specific patient population, your local documentation habits, or rare edge cases unique to your practice. Synthea is the safe first gate, not the final word. A tool that passes synthetic testing still needs a careful, monitored trial on real data before it goes fully live. This staged progression, from synthetic to a limited real trial to full deployment, is the same discipline we apply whether the deployment is cloud-based or kept entirely local.
It does not make a bad AI tool good. Synthea helps you discover the truth about a tool safely. If the tool is poor, Synthea will help you find that out before it costs you anything. That is a feature, not a shortcoming.
How a Practice Should Actually Use This
You do not need to download anything or learn any software. The practical way a practice benefits from Synthea is by making it a standard requirement in how you evaluate any AI vendor. Three simple moves put you ahead of the vast majority of practices:
Ask every AI vendor how they test. When a company pitches you an AI tool, ask a direct question: "Before we involve any real patient data, can you demonstrate this working on synthetic data such as Synthea?" A serious vendor will say yes without hesitation, because responsible ones already test this way. A vendor that insists it needs your live records immediately, before proving anything, is telling you something important about how they work.
Insist on a synthetic-first trial. Make it a condition that any new tool is first exercised against realistic synthetic patients in a test environment, and that you get to see how it performs, before it touches a real chart. This costs you nothing and protects you from the exact scenario that produces breach headlines.
Use a partner who works this way by default. The right technical partner treats synthetic testing not as an extra step but as the obvious starting point. When we build or evaluate AI for a healthcare client, generating a synthetic population and proving the system against it first is simply how the work begins. It is not a premium add-on; it is basic professional discipline.
Open Source vs. Commercial: What Are the Paid Alternatives?
Synthea is not the only way to get synthetic patient data, and it is worth knowing where it sits in the market. Commercial synthetic data platforms exist, chiefly MDClone and Syntegra, both well-funded and both counting major research institutions among their partners, plus Syntho, which works with health systems such as Cedars-Sinai. The broader synthetic data market, healthcare included, is projected to grow from roughly $350 million in 2023 to more than $2 billion by the end of the decade.
The key difference is not price, it is method. MDClone and similar platforms generate synthetic data by learning patterns from your organisation's own real patient data, which produces records that mirror your specific population but means real PHI touches the generation pipeline somewhere and needs its own governance. Synthea generates patients purely from published clinical models and population statistics, with no real patient data involved anywhere in the process. For a practice whose goal is simply testing an AI tool without any privacy exposure, that distinction matters more than any feature list.
Should you build this in-house? Writing a realistic patient generator from scratch is a multi-month undertaking even for a well-resourced team, since it means correctly encoding clinical guidelines, disease progression, and realistic care pathways. Synthea has already done that work and given it away for free. Unless your practice has an unusual need that specifically calls for data mirroring your exact patient mix, a commercial platform like MDClone, there is little reason to build this yourself or pay for one just to safely evaluate a vendor's AI tool.
Frequently Asked Questions
What is Synthea in simple terms?
Synthea is a free, open-source tool made by the non-profit research organisation MITRE. It generates complete, realistic medical records for patients who do not actually exist. Because these patients are invented, their records are not protected health information, so you can use the data to test AI tools, apps, and systems without any privacy risk. It models more than 100 diseases and produces data in the same formats real healthcare systems use.
Is Synthea data safe to use under HIPAA?
Yes. Data generated by Synthea is not protected health information, because it does not describe any real person and cannot be traced back to one. HIPAA's Privacy and Security Rules govern real patient data, so they do not restrict how you use Synthea's synthetic records. The one thing to remember is that HIPAA obligations return the moment you move from synthetic testing to using real patient data, so synthetic data makes the early testing stage free of HIPAA burden, not your whole operation.
Is Synthea really free?
Yes, completely. Synthea is open source and costs nothing to use. There is no licence fee and no per-record charge. The only costs associated with it are the time of the technical person who runs it and loads the data into your test systems, and those are minimal compared with the cost of a single data breach.
Can my practice use Synthea directly, or do we need help?
Synthea is a developer tool rather than a consumer app, so most practices do not run it themselves. In practice, a technical partner generates the synthetic data for you and loads it into a test copy of your systems. You get all the benefit, a safe way to evaluate AI, without needing any technical skill yourself. The more useful thing a practice owner can do is insist that any AI vendor proves their tool on synthetic data before touching real records.
How is Synthea different from just removing names from real patient records?
Removing names from real records, known as de-identification, still starts with real people's data, and there is always some risk that individuals can be re-identified, especially in small or unusual populations. Synthea never starts with a real patient at all. Every synthetic record is built from clinical models and population statistics, so there is no real individual hiding behind it to re-identify. That makes it fundamentally safer than de-identified real data for testing purposes.
Why does synthetic data matter more for AI than for other software?
Ordinary software behaves predictably, so it is easy to test. AI tools built on large language models are probabilistic, meaning the same input can produce different results and a confident-sounding answer can still be wrong. To trust an AI tool, you have to watch it handle a wide range of realistic cases, including difficult ones, which requires a large and varied dataset. Synthetic data lets you do that testing thoroughly without ever exposing a real patient during the phase when you are still deciding whether the tool is any good.
Does Synthea work with our existing EHR?
Very likely, if your systems follow common healthcare standards. Synthea exports data in HL7 FHIR and C-CDA, the standard formats modern EHRs use to store and exchange records, as well as plain CSV. Standards-based and open systems, including open-source platforms like OpenEMR, accept Synthea data readily. A technical partner can confirm compatibility with your specific setup.
The Bottom Line for Your Practice
The reason many practices stall on AI is not that the tools are bad. It is that testing them safely feels impossible, because proper testing seems to demand real patient data, and real patient data is exactly what you cannot afford to put at risk with an unproven tool.
Synthea dissolves that dilemma. It gives you an unlimited supply of realistic, standards-compliant patient records that carry no privacy weight whatsoever, so you can put any AI tool through its paces as thoroughly as you like before a single real chart is ever involved. The largest health agencies in the country already build on this foundation. Your practice can adopt the same discipline.
The practical action is small but powerful: make synthetic-first testing a non-negotiable requirement for any AI tool you consider, and work with a partner who treats it as the natural starting point rather than an afterthought.
If you are evaluating AI for your practice and want to do it the safe way, from synthetic testing through to a carefully monitored real-data trial, ValueStreamAI's healthcare team can guide the whole process. Start with our AI for Medical Practices hub for the full picture, see how a real private deployment came together in our OpenMed medical AI case study, or get in touch to talk through your specific situation.
The tools to test AI safely are free and proven. There is no longer any reason to gamble with real patient data to find out whether AI can help your practice.
ValueStreamAI builds custom agentic AI systems for SMBs and enterprises across the US and UK. Learn more about us →
