stop trying to fix prompts and compare scores on a weekly schedule. That entire approach is a trap, and it's going to waste your time.
your first instinct might be to tweak the AI prompt and then check the leaderboard every week. don't. I've watched teams do this and they end up chasing ghosts. The output from these systems is non-deterministic. run the same query five times in a row, and you'll get five different scores for no apparent reason. at a weekly check-in, you're not seeing a real trend, you're just looking at random variance and inventing a story to explain it.
we got burned by this for weeks before we changed two things.
first, stop looking at the composite score as the main signal. that number is too noisy to be useful at short intervals. instead, track the discrete, binary events. did a new competitor show up in the answer? did your brand get mentioned at all? Did a specific source get cited? those are concrete facts that don't jump around run-to-run. a score changing by 4 points is meaningless. your brand vanishing from the recommended list overnight is a clear, actionable event.
second, establish a noise floor before you even look at the data. we decided nothing matters unless it moved by at least six points. anything smaller gets logged for the trendline but we don't react to it or call it a "change." It's just system jitter.
and a quick point on logging: there's a massive difference between your site being cited as a source and your brand being named as the answer. They are not the same thing and they behave independently. we had a domain cited dozens of times in a single run and still get almost never named as the actual solution. track both.
save yourself the headache and stop looking weekly. monthly cadence forces you to identify real patterns instead of overreacting to noise