METHOD / V1
What ModelMoods measures
ModelMoods is a live summary of how people say the ChatGPT product compares with their own usual experience. It is a human-experience signal—not an intelligence test and not an uptime monitor.
The rating
People rate how ChatGPT felt during the last ten minutes compared with their usual experience: much worse, worse, normal, better, or much better. The stored values −2, −1, 0, +1, and +2 are normalized to a symmetric −1, −0.5, 0, +0.5, and +1 scale for calculation. The headline Mood is centred on Normal, not a one-sided percentage.
Blind before exposed
A report made before the visitor sees the current crowd result is blind. A report made after the result has been revealed is exposed. Both classifications remain supported, but only reports that are both valid and blind affect the official Mood. This reduces—without eliminating—the tendency to follow the crowd.
One person, one current opinion
The server places reports into fixed ten-minute buckets. One anonymous participant has one effective report per product and bucket; another submission in that bucket updates the existing report. Within the selected live window, only each participant’s newest valid blind report counts. The live sample therefore represents unique participants’ latest current opinions in that window.
Live window
ModelMoods tests rolling windows of 10, 30, and 60 minutes, then 3, 6, and 24 hours. It selects the shortest window containing at least 25 unique valid blind participants. If none reaches that threshold, it uses the 24-hour window as the fallback, keeps confidence limited, and lets the prior pull a small sample toward its expected level.
Baseline and smoothing
The baseline covers the 28 days immediately before the selected live window. It uses valid blind observations and becomes ready only with at least 300 observations across seven distinct calendar days. The ten-minute storage rule still limits a participant to one effective report per bucket; participation in separate historical buckets can contribute separate observations.
Before the baseline is ready, the prior Mood is 0, or Normal. Once ready, the prior becomes the historical signed mean. The same baseline also publishes historical worse and better shares. Current reports are then smoothed toward the applicable prior, and the public Mood delta is the smoothed current Mood minus that prior.
Mood states
Classification is symmetric and ordered so the strong states are tested first. If a strong-state threshold is crossed without its statistical guardrail, the reading falls back to the corresponding directional state.
- Degraded
- Delta ≤ −0.25 and the strong negative guardrail passes.
- Below normal
- Not Degraded, and delta ≤ −0.10—including stronger negative deltas whose guardrail does not pass.
- Normal
- −0.10 < delta < +0.10.
- Above normal
- Not Exceptional, and delta ≥ +0.10—including stronger positive deltas whose guardrail does not pass.
- Exceptional
- Delta ≥ +0.25 and the strong positive guardrail passes.
Worse, better, and divided reports
The worse share is the proportion choosing much worse or worse. The better share is the proportion choosing better or much better. Both use the same valid blind live sample as the Mood and remain supporting facts rather than replacements for it.
With at least 25 reports, ModelMoods flags “Reports are divided.” when both shares are at least 25%. The flag prevents opposing experiences from disappearing through cancellation, but it does not replace or suppress the directional Mood.
Confidence
Confidence is separate from direction: fewer than 25 participants is limited; 25–49 is low; 50–149 is medium; and 150 or more is high. These categories describe live sample size, not certainty that a product or model changed.
Technical details
The smoothed Mood is (Σ normalized current ratings + k × prior Mood) / (n + k). Here, n is the deduplicated live sample and k, the configurable prior strength, defaults to 25. Directional deltas default to ±0.10 and strong-state deltas to ±0.25.
For Degraded and Exceptional, ModelMoods calculates a two-sided 95% normal-approximation interval for the raw mean of the bounded −1 to +1 current ratings: mean ± 1.96 × standard error, clamped to that scale. At least 25 live participants are required, and the full interval must lie below the prior Mood for Degraded or above it for Exceptional. The prior strength, deltas, interval multiplier, baseline minima, participant minimum, and divided-share threshold are configurable deployment parameters; the values stated here are the production defaults.
Limits
This is a voluntary sample, not a representative survey. People who are annoyed may be more motivated to report; distribution channels may attract particular users; anonymous identifiers cannot prevent every coordinated manipulation; and people define “usual” differently. The signal cannot identify a root cause, distinguish every ChatGPT model or plan, or prove that an underlying model changed. It reports product experience from participating humans—useful, but not omniscient.