This paper presents a unified framework for natural-language control of robotic manipulators based on the Model Context Protocol (MCP). The system integrates a large language model (LLM) with real-time perception, spatial reasoning, and robot execution, enabling users to command robots through unconstrained natural-language instructions. High-level requests are interpreted by an LLM with structured tool-calling capabilities and translated into executable actions provided by a modular set of MCP tools. The tools interface with a real-time environment layer that manages perception, world modelling, and manipulation, while platform-agnostic controllers enable deployment on multiple robot arms and support multimodal interaction through graphical user interface (GUI), speech, and command line interfaces. Using a NIRYO Ned2 robot arm we evaluate the system on a diverse set of manipulation tasks requiring object identification, spatial reasoning, and multi-step action execution. Experiments demonstrate that the approach achieves reliable task completion despite challenges such as ambiguous object references and visually similar objects. The results highlight the feasibility of combining LLM-based reasoning with classical perception and control for robust, language-driven manipulation. All tools, controllers, and environment components are made publicly available at https: //github. com/dgaida/robotₘcp.
Gaida et al. (2026) studied this question.