benchmark.
Listed 3Updated Sep 11, 2026
Sorted by Most popular
- OSWorld[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environmentsapache-2.0 · linux3.1k
- OSWorld-V2OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasksapache-2.0 · linux303
- XLingEvalCode and Resources for the paper, "Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries"apache-2.0 · linux202