Gut Microbiome Testing: Current Limits

Spread the love

Send a stool sample to a company, get back a report naming hundreds of bacteria with percentages beside each one. The precision of the gut microbiome test is impressive.

Researchers at the US National Institute of Standards and Technology decided to check that precision. They sent identical sample material to seven testing services and compared what came back.

The results reframe what these reports can be read as.

 

Two different technologies, two different answers

Before the study, one distinction matters.

16S rRNA amplicon sequencing reads one marker gene. It identifies microbial communities at a broad taxonomic level but cannot resolve species-level or functional differences with high precision[4]. It’s the cheaper method, and most consumer-facing tests use it[4].

Whole-genome shotgun sequencing reads everything present, giving species-level resolution and functional gene data[4]. More informative, more expensive.

These are not two roads to the same destination. The different techniques can produce different results from the same sample[3].

Beyond sequencing method, providers differ in collection mechanism, DNA extraction protocol, primers, and bioinformatic pipeline[3][4]. Each of those is a place where numbers can diverge.

 

Same material, seven companies

Here’s the experiment.

NIST used a homogenised, well-characterised human stool reference material — deliberately, so that biological variation between donors couldn’t explain any differences[5]. They ordered three kits from each of seven providers and filled them all with the identical material. The companies weren’t told[5].

The headline finding:

Variability between providers was on the same scale as biological variability between different donors[1].

Read that again. The spread between companies analysing *the same stool* was comparable to the spread you’d expect between *different people’s* stool.

The authors attributed this to methodological variability and insufficient quality control[1]. Notably, reproducibility within a single company’s locked-down workflow tended to be very good[1] — the problem is across providers, not necessarily within one.

For five clinically relevant genera — *Bacteroides*, *Bifidobacterium*, *Clostridium*, *Roseburia* and *Faecalibacterium* — reported abundances varied widely, and the comparison ranges each company used differed too[3].

Same stool, two verdicts

The discrepancies didn’t stay technical. They reached the consumer-facing conclusions.

In one case, replicate samples from the same stool were classified as both “healthy” and “unhealthy” — with divergent dietary and functional recommendations[3].

Same material. Opposite verdict. Different advice.

There’s a quality control detail worth adding. One company reported a replicate with a markedly different microbial profile from the same sample, including more than 50% unidentified reads — and still issued a consumer report[3].

More than half the reads unidentified, and a report went out anyway.

 

Nobody agrees what “healthy” means

Even with perfect measurement, a second problem would remain.

There is no consensus on what constitutes a healthy microbiome[3]. Companies varied in how they defined it and in which reference populations they compared against[3].

No universally validated “healthy” profile exists to benchmark an individual result against[4].

This is why the French Society of Microbiology has advised against microbiome testing, citing insufficient knowledge and the personalised nature of defining a healthy microbiome[1][5].

It’s also worth knowing the regulatory position. There are currently no regulatory-approved clinical microbiome diagnostic tests in the US, and only one sequencing-based test carries the CE-IVD designation in Europe[1].

A bit more detail — on what certification actually certifies. Some labs hold CLIA certification, which ensures basic quality standards for laboratory processes. CLIA does not validate the clinical meaning of results; it certifies the process, not the interpretation[4]. Separately, the Microbiota International Clinical Society evaluates these tests against three pillars — analytical validity, clinical validity, and clinical utility — and fully meeting all three remains difficult given the intrinsic variability of microbial composition[4]. Before a result can be considered actionable for dietary guidance, it would need to demonstrate that a specific intervention reliably produces a measurable health outcome[4].

 

Where these tests do have a place

None of this makes the underlying science worthless. It’s worth being precise about the distinction.

Research use is real. Microbiome differences between health and disease states have been documented repeatedly. In inflammatory bowel disease, gut microbiota is consistently less diverse than in healthy controls, and machine learning models have distinguished ulcerative colitis cases from controls with area-under-the-curve values of 0.83 to 0.92[6].

Specific clinical questions may warrant shotgun sequencing — inflammatory bowel disease evaluation, antibiotic resistance gene surveillance, post-transplant monitoring[2].

What doesn’t follow is that a consumer report can tell an individual what to eat. Group-level differences and individual-level prediction are different problems, and we’ve now seen that pattern several times in this series.

 

If you’re in Korea

Two practical notes.

These services are available here, and the same analytical concerns apply — the NIST findings were about methodology, not geography.

Check what a report is being used for. A microbiome report that prompts you to see a doctor about persistent symptoms is being used sensibly. One that prompts you to buy a specific supplement recommended by the same company is a different proposition.

If you have ongoing digestive symptoms, that’s a reason for medical assessment rather than a sequencing kit. Bleeding, unexplained weight loss, persistent changes in bowel habit or symptoms that wake you at night all warrant a clinician rather than a test you order yourself.

 

Closing

I had assumed the limitation here was interpretive — that the measurement was solid and the science just hadn’t caught up on what the numbers mean.

The NIST data says otherwise. The measurement itself diverges between providers by as much as it diverges between different people. And underneath that sits the second problem: no agreed reference for what a healthy result would look like.

A report full of percentages reads like precision. In this case the precision is in the presentation rather than the underlying measurement — and that gap is worth knowing before paying for one.

At a Glance

  • Most consumer tests use 16S sequencing, which cannot resolve species-level or functional detail; shotgun sequencing can but costs more
  • NIST sent identical stool reference material to seven providers, blinded, three kits each
  • Variability between providers matched biological variability between different people
  • Reproducibility within a single company’s workflow tended to be good — the problem is across providers
  • Replicates of the same stool were classified as both “healthy” and “unhealthy”, with different dietary advice
  • One company issued a report on a replicate with over 50% unidentified reads
  • No consensus definition of a healthy microbiome exists; companies use different reference populations
  • No regulatory-approved clinical microbiome diagnostic exists in the US; one CE-IVD test in Europe
  • CLIA certifies the laboratory process, not the clinical meaning of results

※ This article discusses the analytical limitations of commercially available tests and is for general information only. It does not evaluate any specific company or product, and it does not replace medical diagnosis or treatment. Persistent digestive symptoms — particularly bleeding, unexplained weight loss, or ongoing changes in bowel habit — warrant assessment by a clinician rather than self-directed testing.

 

References

  1. “Evaluating the analytical performance of direct-to-consumer gut microbiome testing services”, Communications Biology, https://www.nature.com/articles/s42003-025-09301-3
  2. “Commercial Microbiome Tests: Unreliable Results” (NIST study design and expert commentary), Medscape, https://www.medscape.com/viewarticle/commercial-microbiome-tests-unreliable-results-2026a100079j
  3. “DTC microbiome tests prove wide variability”, AGA Journals News, https://news.gastro.org/issues/2026/april-2026/dtc-microbiome-tests-prove-wide-variability/
  4. “Microbiome Testing & Personalized Nutrition” (sequencing methods, CLIA scope, MICS three pillars), Top Doctor Magazine, https://topdoctormagazine.com/news/microbiome-testing-personalized-nutrition/
  5. “Evaluating the Analytical Performance of Direct-to-Consumer Gut Microbiome Testing Services”, bioRxiv preprint (full methods), https://www.biorxiv.org/content/10.1101/2024.06.05.596628.full.pdf
  6. “16S rRNA and metagenomic shotgun sequencing data revealed consistent patterns of gut microbiome signature in pediatric ulcerative colitis”, PMC, https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9018687/

Leave a Comment