munch2u-a11y / FP-AMB Star 4 Code Issues Pull requests First-Person Agent Memory Bench. 10 Categories including fact recall, multi-hop links, temporal reasoning, fact overwrites, speaker traps, refusal, credibility, and agentic tool usage. 540K token / 60 session corpus, all in first person. Dynamic output-answer-key portion. Comprehensive report with visuals and miss breakdown. ai-safety memory-benchmark ai-agent ai-governance ai-system-design reasoning-benchmark longmemeval ai-agent-memory ai-memory-system ai-agent-testing tool-use-behavior-identification beam-benchmark ai-agent-benchmark locomo-benchmark ai-memory-management context-memory-benchmark Updated Aug 27, 2026 Python