From Fuzozo to Pophie: What AI Companion Devices Teach Us About Building the Next Generation of IoT Products
- Last Updated: August 21, 2026
Lawrence Wu
- Last Updated: August 21, 2026



For decades, connected devices have been built around commands. A user presses a button. A sensor reports a reading. A mobile app sends an instruction. Even voice assistants largely followed the same pattern by waiting for a wake word before processing a request. A new generation of AI-native devices is changing that interaction model.
Products like AI companions, consumer robots, assistive technologies, educational devices, and wellness products are expected to hold natural conversations, remember previous interactions, recognize different users, respond emotionally, and stay available throughout the day. Instead of reacting to isolated commands, they participate in continuous interactions.
That seemingly small shift fundamentally changes the engineering requirements behind the device. Building these experiences is no longer just about integrating a language model. Developers must solve a new class of real-time systems problems involving latency, speech recognition, interruption handling, identity, memory, synchronization, and global infrastructure.
Two recent products illustrate this particularly well: Robopoet's Fuzozo and InsBotics' Pophie. Although they target different audiences, they reveal a broader lesson about where IoT architecture is heading.
When users are speaking naturally, every pause becomes noticeable. Conversations involve overlapping speech, interruptions, changing speakers, emotional tone, background noise, and long-running context. Unlike issuing a command to a device, conversation is continuous and highly sensitive to timing.
The engineering challenge shifts from transmitting data efficiently to maintaining a believable interaction. Instead of optimizing only network throughput or cloud connectivity, developers must optimize something much harder: the flow of human conversation.
Most engineers understand that lower latency improves responsiveness. For conversational devices, latency affects something deeper. It shapes whether an interaction feels natural. People instinctively expect conversations to flow without awkward pauses. Delays that might be perfectly acceptable for a dashboard update become distracting during spoken dialogue. Users begin talking over the device, repeat themselves, or disengage entirely.
This challenge became particularly important for Fuzozo, Robopoet's AI-powered plush companion. Unlike a smart speaker sitting on a kitchen counter, Fuzozo is designed to feel like an emotionally engaging companion rather than a gadget. Conversations needed to remain immediate, interruption-friendly, and emotionally responsive despite the constraints of a compact plush device. That required solving multiple engineering problems simultaneously:
Rather than optimizing only inference speed, the system needed low-latency speech recognition, voice activity detection (VAD), interruption handling, and conversational orchestration working together.
The results suggest these interaction improvements had a measurable effect on user engagement. Within 10 months of its launch in 2025, Robopoet sold more than 250,000 Fuzozo units, while users averaged nearly 50 minutes of conversation each day. The company also generated more than 1,000 JD.com pre-orders within the first 10 minutes of launch, highlighting strong early consumer demand.
Those numbers reflect more than product popularity. Sustaining nearly an hour of daily conversation suggests the underlying interaction experience remains compelling long after the novelty wears off.
Many connected devices only need to recognize that someone is present. AI companions need to understand who is present. This distinction becomes increasingly important as conversational devices move into shared households, classrooms, healthcare environments, and hospitality settings. Without user identity, conversations quickly lose continuity. Personal preferences disappear. Memories become inconsistent. Context is repeatedly reset.
For Fuzozo, voiceprint recognition became an important architectural capability rather than an optional personalization feature. Instead of treating every conversation as new, the device can recognize individual users and personalize future interactions based on previous conversations and preferences. This enables several capabilities that become increasingly valuable over time:
As AI-native devices become shared household products rather than personal gadgets, identity management will likely become a standard component of conversational infrastructure.
Speech alone no longer defines conversational AI. Many AI-native products now combine voice, vision, memory, movement, and emotional expression into a single interaction.
This creates an entirely different synchronization challenge. Unlike traditional IoT systems, multiple outputs must remain coordinated in real time. If speech arrives immediately but facial expressions, eye movements, or robotic gestures lag behind, the interaction quickly feels artificial.
This challenge is evident in Pophie, the emotionally intelligent AI companion developed by Singapore-based InsBotics. Pophie combines conversational AI with expressive robotics, contextual memory, facial recognition, voiceprint recognition, and natural movement to create interactions that feel emotionally present throughout the day. Achieving that experience required synchronizing several systems simultaneously:
Rather than treating robotics and conversation as separate systems, they operate as one coordinated interaction pipeline. This architecture contributed to strong early adoption.
Following launch, Pophie attracted 523 Kickstarter backers while raising US$170,880 within its first 12 hours. Active users averaged 185 daily interactions, while the device maintained an always-on presence exceeding 12 hours per day.
These usage patterns highlight an important trend: As AI companions become more embedded in everyday life, developers are no longer optimizing for short commands. They are optimizing for persistent presence.
Large language models continue improving rapidly, and increasingly similar capabilities are becoming available across the industry.
What differentiates AI-native devices is often no longer the model itself. Instead, differentiation increasingly comes from the infrastructure surrounding it. Real-world conversational experiences depend on many systems operating together:
Any weakness across these layers directly affects the overall user experience. Developers building AI-native hardware are therefore spending less time asking, "Which model should we use?" and more time asking, "How do we make conversations feel natural under real-world conditions?"
The IoT industry has spent years connecting devices to networks. The next phase will focus on making those devices genuinely interactive. Whether the product is a companion robot, educational assistant, healthcare device, or smart home system, users increasingly expect technology to listen naturally, remember context, recognize individuals, and respond in real time.
Fuzozo and Pophie demonstrate that delivering these experiences requires far more than embedding an AI model into hardware. It requires building an interaction stack where speech recognition, latency, synchronization, identity, memory, and conversational orchestration work together as a unified system. As Physical AI continues expanding into consumer devices, robotics, and everyday environments, conversation will no longer be viewed as an application feature. It will become foundational infrastructure for the next generation of IoT products.
The Most Comprehensive IoT Newsletter for Enterprises
Showcasing the highest-quality content, resources, news, and insights from the world of the Internet of Things. Subscribe to remain informed and up-to-date.
New Podcast Episode

Related Articles
August 21, 2026

August 19, 2026

August 13, 2026
