Hacker News new | ask | show | jobs
user: brian8620
created: 2026-06-03
karma: 6

I'm building Veritrooper — it audits how accurately an LLM answers from a body of written rules (tax code, safety regs, drug labels, filings… anything written) and hands back a per-question record of where and why it's wrong. It started as an AI Dungeon Master that kept confidently breaking D&D's own rules. That failure mode — sounding certain about something the source doesn't actually support — turned out to be the real problem with LLMs on regulated text, so I went and built the thing that catches it. The part I care about: it isn't just an LLM grading an LLM. The clear-cut calls are deterministic; only the genuinely contested ones go to a model from a different vendor, and nothing ever grades its own vendor. Every number reproduces from the logs, and the run records are downloadable — so you can check my work. Solo dev, first real software project. veritrooper.com · contact@veritrooper.com

submissions:

Show HN: Veritrooper – find what your AI gets wrong about your own docs
2 points | 0 comments
0 points | 0 comments
0 points | 0 comments
Show HN: Audit any AI/data pairing with Veritrooper
2 points | 1 comments
0 points | 0 comments
0 points | 0 comments
0 points | 0 comments
0 points | 0 comments