نيورا

Benchmark

LongMemEval · Hit@10
Recall accuracy
79%
+37.5 pts vs baseline
Baseline recall
41.5%
embedding-only
Relative gain
1.9×
near 2× recall
End-to-end accuracy
84.2%
full QA · gpt-4.1 judge · 2026-07
Recall across runs

Memory retrieval trend

70
80
90
41.5
79
baselineneura
Run config

Suite

Probe setLongMemEval
MetricHit@10
Baseline41.5% (embedding-only)
MethodBM25 + embeddings + weighted RRF + temporal re-ranking
Measured2026-06
End-to-end84.2% answer accuracy (2026-07)
Emotion vs human self-reports

Does the AI feel what people reported feeling?

Emotion category
59.6%
vs 7.7% chance (13 emotions)
Felt sentiment
ρ 0.64
correlation with the person’s own rating
Ground truth
n=52
crowd-enVENT (human self-reports) · 2026-07

Hit@10 measures how often the right memory is surfaced in the top 10 recalled items across long, multi-session conversations. Neura nearly doubled recall over the embedding-only baseline.

End-to-end accuracy: the full pipeline (ingest a year of conversations, retrieve, answer) graded question-by-question by the standard judge. Direct single-session recall scored 95.7-100%.

Emotion validation: events lived and self-rated by real people were shown to the appraisal system. It named the right emotion family 59.6% of the time (13 emotions, 7.7% chance) and its felt positivity/negativity correlated with the person’s own rating at rho = 0.64 (n = 52, stratified across all 13 emotions).

مُقاس على LongMemEval (2026-06). الطريقة: BM25 + embeddings + weighted RRF + temporal re-ranking.

سياسة الخصوصيةشروط الاستخدام