arxiv:2609.16251
Published on Sep 14
· Submitted by
Zihan on Sep 21
Upvote
5
Authors:
,
,
,
,
,
,
Abstract
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, structured artifacts. Mechanical computer-aided design (CAD) is a particularly demanding setting: an agent must manipulate geometry and constraints over long interaction horizons while producing a native project whose dimensions, construction structure, and downstream engineering state remain valid. We introduce CADWorld, a benchmark for long-horizon computer use in FreeCAD. CADWorld contains 200 tasks spanning 11 mechanical-CAD workflow categories, including sketching, part modeling, assembly, CAM, FEM, measurement, mesh processing, and technical drawing. Agents operate through screenshots and GUI actions, while success is determined by task-specific executable checks over saved FreeCAD artifacts and auxiliary outputs, covering geometric properties, parametric structure, constraints, manufacturing state, and simulation results. Across seven current agents on the full benchmark, the strongest agent achieves 17.5% success, compared with an 87.0% expert reference pass. We find that weaker agents often fail before producing a valid artifact, whereas stronger agents increasingly fail on structural, geometric, and construction-process requirements. CADWorld therefore exposes a gap between general GUI competence and reliable execution of persistent, verifiable engineering workflows. Project accessible at https://cad-world.github.io.
View arXiv page View PDF Add to collection
Community
Paper submitter about 24 hours ago
Been recommended by NielsRogge to submit here
about 13 hours ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
-
RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents (2026)
-
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions? (2026)
-
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software (2026)
-
Qwen-CUA: Native Computer Use for (almost) Everything (2026)
-
ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions (2026)
-
LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows (2026)
-
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on HF Mirror checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
· Sign up or log in to comment
Upvote
5
Get this paper in your agent:
hf papers read 2609.16251
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.16251 in a model README.md to link it from this page.
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.16251 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.