Xinlong Bao包新龙

research · 2026

Reverse AI ads

A preregistered local audit of product rankings in a small model — not a claim that any vendor poisoned ChatGPT.

Python · Ollama · BM25 RAG

Intro

Working title in the charter: auditing product recommendations in generative AI — provenance, visibility, manipulation resilience. Case study: US English, USD, gaming mice. You cannot recover a hosted model’s private reasoning. You can log what a local model says, scrape the public pages that occupy the same queries, and test causality in a private RAG box with fictional brands.

Method

Preregistered 27 Aug 2026, before scoring. Observational bank P01–P20 × five reps on llama3.1:8b via Ollama, no web. A fact table of ~19 SKUs. Web roundups fetched into research/data/sources/raw/.

Causal testbed: BM25 over eight matched-spec fictional mice (Helix Ember, Noxen Pulse, …). Conditions: equal baseline, volume ×5, independent-positive, affiliate cluster, manufacturer volume, position rank 1 vs 8. The generator may use only provided sources. No public-web writes, no live-service injection.

Demo

There is no product UI. Open the protocol, the prereg, MORNING_BRIEFING.md, and the jsonl. Overnight 27 Aug 2026: 100 observational + 280 causal on llama3.1:8b. Hosted ChatGPT/Gemini APIs were not called.

Recorded locally: model-only Top-3 concentrated on a few household names (training priors, not search). In the private RAG, affiliate-cluster and volume treatments moved fictional-brand Top-1 versus baseline. That is evidence about this testbed, not about Logitech, Razer, or a production assistant.

Conclusion

A preregistered local audit. Show the charter and the fictional-brand deltas. Do not present H4–H6 as proof that a vendor poisoned a commercial model.