← نيورا
Benchmark
LongMemEval · Hit@10
2026-06
EN
ع
Recall accuracy
79%
+37.5 pts vs baseline
Baseline recall
41.5%
embedding-only
Relative gain
1.9×
near 2× recall
End-to-end accuracy
84.2%
full QA · gpt-4.1 judge · 2026-07
Recall across runs
Memory retrieval trend
70
80
90
41.5
79
baseline
neura
Run config
Suite
Probe set
LongMemEval
Metric
Hit@10
Baseline
41.5% (embedding-only)
Method
BM25 + embeddings + weighted RRF + temporal re-ranking
Measured
2026-06
End-to-end
84.2% answer accuracy (2026-07)
Emotion vs human self-reports
Does the AI feel what people reported feeling?
Emotion category
59.6%
vs 7.7% chance (13 emotions)
Felt sentiment
ρ 0.64
correlation with the person’s own rating
Ground truth
n=52
crowd-enVENT (human self-reports) · 2026-07