Benchmark evaluation demonstrates improved scientific data visualization across diverse large language models, highlighting the utility of visual feedback mechanisms.
Key Points
Automated scientific data visualization performance improves across diverse large language models when guided by the model-agnostic MatPlotAgent agentic framework.
Across 100 human-verified test cases in the MatPlotBench benchmark, visual feedback and iterative debugging enabled substantial gains in code generation accuracy.
Assessment using the MatPlotBench framework and GPT-4V evaluation correlates strongly with human annotations, highlighting viable automated multi-modal LLM scoring.