Loading
A benchmark for evaluating scientific data analysis and visualization agents, addressing the lack of principled and reproducible benchmarks for agentic systems in scientific visualization tasks, with potential impact on the development of more effective and efficient SciVis agents
“arXiv:2603.29139v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid pro…”
Read the source →STATUS
ACTIVE
CATEGORY
Models
EVIDENCE
Not yet assessed
ENTITY
SciVisAgentBench, Large Language Models (LLMs), arXiv
DECISION
Automated · no editorial override
LAST OBSERVED
Jul 21, 2026