Hi Erik and Scott,
Following our email discussion, I am opening this issue to document the terminal-drift sampling behaviour we encountered in mixed Gaussian–Bernoulli models and the conditional non-centering that substantially improved it.
Motivation
Our application contains Gaussian traits together with a Bernoulli-logit trait. During diagnosis, the sampling problem could already be reproduced in a balanced 190-tip E+P model with all cross-effects fixed to zero and correlated drift disabled.
The corresponding univariate processes sampled well independently, whereas the mixed E+P model did not. The main failure was concentrated in the Bernoulli diffusion scale Q_sigma[P].
Minimal diagnostic
For the same balanced 190-tip dataset:
- P only:
Q_sigma[P] R-hat = 1.001, bulk ESS = 1,359
- centred E+P: R-hat = 1.499, bulk ESS = 7.7, tail ESS = 18.4, E-BFMI = 0.064–0.190
With cross-effects and correlated drift disabled, we verified numerically that the E+P target factorises into the corresponding univariate targets up to a constant.
We then noted that the P-only model represents the terminal innovation on a standard-normal scale, whereas the mixed model uses the realised-scale Bernoulli terminal drift.
Conditional non-centering
We constructed a target-equivalent diagnostic version in which the Bernoulli terminal innovation was conditionally non-centred while retaining the Gaussian dimensions and the Cholesky covariance structure.
Target-density and gradient equivalence were checked before sampling.
For the same E+P model:
- conditionally non-centred E+P:
Q_sigma[P] R-hat = 1.002, bulk ESS = 1,281, tail ESS = 1,010, E-BFMI = 0.839–1.024
There were no divergences or maximum-treedepth hits, and all monitored population/process parameters passed our convergence criteria.
We also applied the corresponding conditional transformation to the full three-trait D0 model. This substantially improved Q_sigma[P], E-BFMI, and the divergence count, but the full model still did not pass the convergence gate; additional geometry remained in several cross-trait/drift parameters.
Possible package-level direction
A possible package-level change would be to allow the Bernoulli terminal innovation to use a conditional non-centred representation in mixed Gaussian–Bernoulli models while retaining the full covariance structure.
As Erik noted in our email discussion, this uses the same conditional-Gaussian algebra already present in the generated-quantities calculation, applied in the reverse conditional direction.
The compact 190-tip diagnostic package has already been shared in our email thread. I can add a minimal public reproducer here if useful for implementation or testing.
Thanks again for taking a look at this.
Hi Erik and Scott,
Following our email discussion, I am opening this issue to document the terminal-drift sampling behaviour we encountered in mixed Gaussian–Bernoulli models and the conditional non-centering that substantially improved it.
Motivation
Our application contains Gaussian traits together with a Bernoulli-logit trait. During diagnosis, the sampling problem could already be reproduced in a balanced 190-tip E+P model with all cross-effects fixed to zero and correlated drift disabled.
The corresponding univariate processes sampled well independently, whereas the mixed E+P model did not. The main failure was concentrated in the Bernoulli diffusion scale
Q_sigma[P].Minimal diagnostic
For the same balanced 190-tip dataset:
Q_sigma[P]R-hat = 1.001, bulk ESS = 1,359With cross-effects and correlated drift disabled, we verified numerically that the E+P target factorises into the corresponding univariate targets up to a constant.
We then noted that the P-only model represents the terminal innovation on a standard-normal scale, whereas the mixed model uses the realised-scale Bernoulli terminal drift.
Conditional non-centering
We constructed a target-equivalent diagnostic version in which the Bernoulli terminal innovation was conditionally non-centred while retaining the Gaussian dimensions and the Cholesky covariance structure.
Target-density and gradient equivalence were checked before sampling.
For the same E+P model:
Q_sigma[P]R-hat = 1.002, bulk ESS = 1,281, tail ESS = 1,010, E-BFMI = 0.839–1.024There were no divergences or maximum-treedepth hits, and all monitored population/process parameters passed our convergence criteria.
We also applied the corresponding conditional transformation to the full three-trait D0 model. This substantially improved
Q_sigma[P], E-BFMI, and the divergence count, but the full model still did not pass the convergence gate; additional geometry remained in several cross-trait/drift parameters.Possible package-level direction
A possible package-level change would be to allow the Bernoulli terminal innovation to use a conditional non-centred representation in mixed Gaussian–Bernoulli models while retaining the full covariance structure.
As Erik noted in our email discussion, this uses the same conditional-Gaussian algebra already present in the generated-quantities calculation, applied in the reverse conditional direction.
The compact 190-tip diagnostic package has already been shared in our email thread. I can add a minimal public reproducer here if useful for implementation or testing.
Thanks again for taking a look at this.