Shengwang R2 Conversational AI Robot Development Kit

Industry Category
Sensing Systems Audio module Speech and Semantics Module
联系二维码

Scan to Inquire About the Product

Product Description

The R2 Conversational AI Robot Development Kit is Shengwang’s integrated solution for desktop robots and emotional companion robots. Building upon the real-time AI voice interaction capabilities of the R1 series—including full-duplex conversation, background noise reduction and intelligent interruption—the R2 kit adds local visual recognition and multi-degree-of-freedom motion control, achieving the critical leap from "hearing and speaking" to "seeing and moving". Product Introduction: The R2 kit integrates powerful NPU and ISP, providing a complete edge-side multimodal AI solution. It enables complex visual functions such as face tracking, gesture recognition and object tracking, combined with multi-degree-of-freedom motion control to allow robots to perform emotionally engaging physical interactions such as "walking up to greet users" and "turning to look at the speaker". Key Features: Multimodal Interaction: Integrates voice, vision and motion control capabilities Emotional Design: Establishes emotional connections through visual gaze and body movements All-Scenario Adaptation: A single base unit supports a wide range of scenarios, including educational companionship, office collaboration, home interaction and wearable recording Rapid Development: Provides a turnkey solution that significantly shortens the time to market Main Application Scenarios: Desktop emotional robots, intelligent learning assistants, meeting assistants, home visual control centres, lightweight AI recorders, etc.

Product Specifications

The R2 All-Scenario AI Robot Development Kit achieves multiple breakthroughs in technical performance. In terms of voice interaction, it fully inherits the industry-leading capabilities of the R1 series, including real-time AI voice interaction technologies such as full-duplex conversation, background noise reduction and seamless interruption handling. Conversation latency can be as low as 650 ms, with interruption response times as low as 340 ms, providing near-human-like conversation response speed and rhythm. In complex environments, it can block 95% of ambient human voices and noise interference, achieving precise recognition of conversational speech. In terms of visual capabilities, by utilising the powerful integrated NPU and ISP of the BK7259 chip, the R2 adds local visual recognition and processing capabilities, supporting functions such as face tracking, gesture recognition and object tracking. Visual processing latency is controlled at the millisecond level, enabling real-time recognition and response to visual commands. In terms of motion control, it supports precise multi-degree-of-freedom control, combined with visual and voice functions to achieve emotionally engaging physical interactions such as "walking up to greet users" and "turning to look at the speaker". The kit utilises a low-power design, supporting an ultra-long standby time and effectively alleviating concerns regarding device battery life. It also supports 47 languages, achieving low-latency responses through servers deployed overseas, alongside real-time multilingual conversion and content output. In terms of development efficiency, it takes just one hour to run a demo and one day to complete product prototype sampling, significantly shortening the product development cycle.

Company Profile

Shanghai Agora Technology Co., Ltd.

Founded in 2014, Agora is the global pioneer in real-time audio and video cloud services, delivering the best possible experience for multimodal real-time interactions between people, between people and agents, and between agents. By simply calling Agora’s APIs, developers can build a wide range of real-time interactive scenarios within their applications, such as conversational AI, audio and video calls, and live streaming. Agora’s APIs have empowered more than 20 industries—including AI, social live streaming, education, gaming, IoT, finance, healthcare and enterprise collaboration—across a total of over 200 use cases. On 26 June 2020, Agora, Inc., the parent company of Agora, successfully listed on the Nasdaq under the ticker symbol “API”. As of 31 December 2025, the number of registered applications worldwide had exceeded one million. Throughout 2025, the service delivered over one trillion minutes of communication. Agora has launched the world’s first conversational AI engine, empowering developers to build real-time voice dialogue experiences based on any large language model. It has also created the world’s first—and to date, largest—real-time audio and video network: the Software-Defined Real-Time Network (SD-RTN™). Agorapulse’s technology services cover more than 200 countries and regions worldwide, with clients including industry giants, unicorns and start-ups such as Xiaomi, Momo, Douyu, Bilibili, Xiaohongshu and Yalla; Agorapulse’s technology is also utilised by renowned global enterprises such as HTC VIVE, The Meet Group and Bunch.
Visitor Registration Call us WeChat

Leave Messages

Exhibitor Application
Visitor Registration Contact Add WeChat

Reserve a Booth

Exhibitor Application

Medium/EN