Evalio

Evaluate.
Compare.
Explore data.
Make better decisions.

Benchmark models, track every evaluation run, and turn raw results into confident choices.

Projects

Name DatasetsModelsUpdated
128May 12
95May 10
73May 8
42May 6
1–4 of 4

LLM Benchmarking

Recent Evaluations

May 12GPT-4o on MMLU72.1
May 11Claude 3.5 on MMLU64.8
May 10Llama 3 70B on GSM8K61.3
May 10Gemini 1.5 Pro on TruthfulQA55.7

Leaderboard · Top 5

1GPT-4o72.1
2Claude 3.564.8
3Llama 3 70B61.3
4Mistral Large58.2
5Gemini 1.5 Pro55.7

Results

DatasetModel ScorePrice / 1K tokUpdated
1–5 of 5

Evaluation Tasks

To Do 3
Run MMLU on Gemini 1.5
ARMay 14
Analyze new dataset
JMMay 15
QA review
ARMay 16
In Progress 2
Run Llama 3 70B on GSM8K
JMMay 12
Analyze results
ARMay 12
Done 4
Run GPT-4o on MMLU
ARMay 12
Run Claude 3.5 on MMLU
JMMay 12
Baseline comparison
ARMay 11
Prompt regression suite
JMMay 10

File Manager

datasetsMay 12
resultsMay 11
benchmark_report.pdfMay 10
pricing_snapshot.csvMay 9
notes.txtMay 8
Trash is empty
1–5 of 5

Notification Center

Evaluation completedGPT-4o on MMLU finished with score 72.1
2m
New comment on “LLM Benchmarking”@alex left a comment: “GSM8K numbers look off, can we re-run?”
1h
Evaluation failedGemini 1.5 Pro on GSM8K hit a rate limit — retrying
3h
You're all caught up — no unread notifications.

Settings

Profile

How you appear across your workspace.

AR
Alex Riveraalex@evalio.io · Admin

Project created!

Multimodal Eval” has been successfully created and is ready for its first evaluation.