![]()

Seoul, South Korea – July 23, 2026 – Spot, the four-legged robot developed by Boston Dynamics and currently exhibited at the Seoul Robot & AI Museum, now responds to Korean voice commands from visitors.
The Korean-language voice control feature has been deployed as a live public service at the museum, allowing visitors to experience it on site. When visitors say commands such as “greet,” “sit,” “praise,” “lie down,” or “stretch” in Korean, Spot recognizes the command and performs the corresponding action.
The Seoul Robot & AI Museum, also known as Seoul RAIM, requested the feature from VIDRAFT, a South Korean artificial intelligence company. VIDRAFT implemented the voice control capability using a third-party on-device AI system. Boston Dynamics was not involved in the development of this feature, and neither Spot’s hardware nor its manufacturer firmware was modified. The function was added through an external device attached to the robot.
Adding Korean interaction to robots built for a global market
Many commercial robots are designed and shipped primarily for English-language operating environments. As a result, institutions in non-English-speaking countries often have to wait for manufacturers to provide local-language support or rely on staff to operate robots manually during exhibitions.
The Seoul RAIM case shows another approach: a museum operator added local-language interaction without waiting for the manufacturer’s development schedule and without modifying the robot itself. The project demonstrates that language accessibility in commercial robotics does not always have to be solved exclusively by the original manufacturer.
Spot is one of the museum’s major exhibits, drawing visitors’ attention as it walks autonomously inside the facility. With Korean voice control added, visitors can do more than watch the robot move. They can speak to it directly and see it respond in real time.
A small on-device AI module mounted on Spot
VIDRAFT’s system uses a compact on-device AI unit mounted on Spot’s back. The device is based on VIDRAFT’s in-house physical AI vision-language model and inference acceleration technology, implemented on an NVIDIA Jetson Nano board. The unit is designed to operate for up to nine hours on its standard battery.
Voice recognition and command interpretation are processed locally on the device attached to the robot. Visitor voice data is not sent to an external cloud server.
This approach offers several practical advantages in a public exhibition setting. It reduces latency by eliminating round trips to a remote server. It lowers the burden of handling personal voice data in a space visited by the general public, including children. It also makes the system less dependent on wireless network conditions, which can fluctuate in crowded museum environments.
Toward open-ended voice conversation
At present, Spot recognizes a predefined set of Korean commands designed for museum operations. Seoul RAIM plans to work with VIDRAFT on a next step: enabling visitors and Spot to engage in two-way spoken interaction.
A Seoul RAIM official said the current system allows visitors to speak predefined Korean commands and see Spot respond. The museum plans to consult with VIDRAFT on adding a function that would allow visitors and Spot to communicate through spoken dialogue in the future.
If implemented, visitors would no longer need to memorize specific commands. Instead, they could ask questions and interact with the robot in more natural language. However, recognizing fixed commands and understanding open-ended conversation are technically different challenges. In a public exhibition environment, such a system must be tested not only for accuracy, but also for safety, inappropriate responses and the ability to say when it does not know the answer.
Statement from VIDRAFT
“Robots understanding human language should not be something only the manufacturer can enable,” said Minsik Kim, chief executive officer of VIDRAFT. “Seoul RAIM asked an important question: Can a robot understand the language of the country it is operating in without modifying the robot itself? This project shows one answer in a public museum setting.”
“Visitors in Chang-dong can now speak to Spot in Korean and see the robot respond,” Kim said. “This approach is not limited to Korean. The same structure can be adapted to other language environments.”
“The next step is to move beyond fixed commands into open-ended conversation,” he added. “But when a robot speaks with people in a public exhibition, it is just as important for the robot to know when it does not know something as it is to answer correctly. We intend to release that capability only after sufficient testing.”
About VIDRAFT
VIDRAFT is a South Korean artificial intelligence company founded in March 2024. The company focuses on proprietary foundation models, quantum computing, physical AI and on-device AI for robotics.
VIDRAFT describes itself as the first Korean company to independently develop and publicly release an open-source foundation model on Hugging Face. In July 2026, the company released its proprietary foundation model, Aether-7B-5Attn, under the Apache-2.0 license. The model is designed as a fully open foundation model, with not only model weights but also a training data recipe, training code, full 160,000-step training logs and intermediate checkpoints.
VIDRAFT is focused on building artificial intelligence systems that operate in real-world environments by combining language models, on-device AI and physical AI.
Media Contact
Company Name: VIDRAFT
Contact Person: Minsik KIM
Email: Send Email
Country: South Korea
Website: http://www.vidraft.net