Writing
Blog
Good science starts with a hunch from an anecdotal observation, a pattern, or just a sense that something different is happening. These essays explore how we turn imperfect observations into reliable insight across biology, medicine, statistics, visualization, and AI.
Four Agents, No Ground Truth
An error-correcting loop for scientific prose: a draft is attacked by deterministic checks and four walled-off LLM judges, its objections tracked across revisions, until it stabilizes or halts for a human. The hard part is not building it — it is deciding what “better” means when there is no ground truth. GitHub ↗
Pregnancy according to My Apple Watch
An Apple Watch recorded my second pregnancy almost from the beginning. On VO₂ max, gait asymmetry, double support time — and what it means when one metric says you've recovered and another says you haven't.
Ethomics 3 - Benchmark Scores Have a Goodhart Problem. So Did My Fly Feeding Data.
Once a metric becomes the target, it starts attracting loopholes. How a problem we ran into in fly neuroscience maps onto AI evaluation, digital health, and anywhere a proxy gets mistaken for the construct.
Ethomics 2 - Whorlmap: An Uncertainty-aware Heatmap
Ordinary heatmaps quietly discard confidence interval information: one color per cell, no memory of how well the effect was estimated. On encoding the full bootstrap distribution inside each cell — and why it matters when you need to compare not just what changed, but how certain you are.
Ethomics 1 - What Is a Hungry Fly?
On behavioral state as a high-dimensional object. How the Espresso platform and DESTRA benchmarking framework distinguish genuine hunger from superficially similar signatures.
Effect Sizes 2 - Interactions, Delta-Deltas, and How We Measure Uncertainty
A delta-delta is the same contrast as a regression interaction term or a two-way ANOVA interaction. But bootstrap confidence intervals diverge from model-based uncertainty — and that difference is the whole point.
DeepLabCut and How I Learned about Pretrained Models
On adapting a pretrained vision model for animal behavior tracking — and what that shift revealed about the difference between a model as object and a model as tool.
Decision at the Odor Port
A novel odor-guided intertemporal choice task for mice, and what serotonin does at the moment of decision. How to make a hidden computation measurable.
Effect Sizes 1 - Getting Over ANOVA
Why effect size estimation replaces significance testing. The logic behind DABEST 2.0: bootstrap confidence intervals, estimation graphics, and the delta-delta for factorial designs.