Y
Hacker News
new
|
ask
|
show
|
jobs
user:
jflynt76
created:
2026-05-12
karma:
8
submissions:
0 points
|
0 comments
A benchmark revealing an average memory-retrieval accuracy of 9%
5 points
|
0 comments
A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
4 points
|
0 comments
Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
4 points
|
0 comments
What happens when OKF runs inside an AI tool
5 points
|
0 comments
0 points
|
0 comments
Why most AI evals would miss the Linear sales email failure
6 points
|
2 comments
Two AI judges scored our agent's answer 0.85, but it never opened the file
6 points
|
0 comments
AI Governance Cannot Be a Tool Call
5 points
|
0 comments
AI Memory Is Still Thinking Like Search
5 points
|
0 comments
Show HN: Synthetic corporate dataset generator for AI agent evaluation
3 points
|
0 comments
Why MCP is the wrong abstraction for memory
4 points
|
0 comments
Show HN: Tenure – Traceable AI memory with configurable memory modes
3 points
|
0 comments
You Don't Need a GitHub Copilot Subscription to Use VS Code AI Features
5 points
|
1 comments
AI Memory Proves Inefficient: Tenure Project Detects 95% Error Rate
5 points
|
0 comments
0 points
|
0 comments