Mynd Healthcare

Report / Paper metadata

LGTM: Gaussian process modulated neural topic modeling for longitudinal microbiome

LGTM summarizes repeated microbiome samples as mixtures of microbial co-abundance patterns and models their association with time and other metadata. Its prediction tests concern microbial relative abundances, not diagnosis, treatment response or causal effects of diet and disease.

What is a microbial topic?

A topic is a model-learned pattern of taxa whose relative abundances tend to vary together. Each sample is represented as a mixture of these topics. The paper explicitly uses this as a statistical grouping, not proof of an ecological guild, shared function or a direct interaction between microbes. Topic proportions are not absolute microbial loads.

How does time enter the model?

An autoencoder learns a compact representation of the observed microbiome, while additive Gaussian-process components describe associations with time, subject identity and other measured covariates. A shared topic-to-taxon mapping turns the representation back into relative abundances. A smooth kernel helps model gradual trends and irregular sampling, but abrupt events that were not recorded or were poorly sampled may be smoothed over.

Which data and prediction tests were used?

The paper analyzes longitudinal child microbiome data from Dhaka and DIABIMMUNE, plus HMP2 data from participants with and without inflammatory bowel disease. Imputation and forecasting tests focused on the two child datasets. Imputation used sample-level splits that included samples from all subjects in training. Forecasting used subject-level splits and supplied the first half of observations for held-out subjects to predict their later samples. These are different tasks and should not be treated as testing the same form of generalization.

What did the benchmarks establish?

The authors report better prediction than the BRITS and SAITS comparison methods under their evaluation setup, and a small accuracy trade-off against DGBFGP in exchange for a more interpretable decoder. Topic-discovery evaluation is a reconstruction task, with the observed microbiome available during testing, rather than forecasting unseen profiles. Topic diversity and repeat-run stability improved, especially with a structured initialization. Neither reconstruction scores nor stability prove that every topic has a biological function.

What do disease and dietary associations mean?

The model linked topic variation to age and feeding in child cohorts, and to individual differences, dysbiosis, disease labels and dietary metadata in HMP2. Its covariate-importance measures summarize variation within the fitted model, not causal effects. Correlated metadata can complicate attribution. Associations with foods or disease labels are not dietary advice, a diagnostic threshold or evidence that changing a food prevents inflammatory bowel disease.

What limits remain?

The model uses relative-abundance profiles and does not explicitly model sample-specific read depth, overdispersion or abundance-level zeros. Topic number is a resolution choice that needs sensitivity checks; initialization affects detailed topic composition. Analysis is currently cohort-specific, so transfer across cohorts needs compatible taxa and metadata plus explicit validation. This original account uses the full paper and keeps future clinical-outcome prediction proposals separate from the microbial-profile tests actually performed.

Kind
Paper metadata
Identifier
PMC13644641

View original record ↗

Worldwide health knowledge / Healthcare records