After a year of building real-time AI apps on Zoom, I've noticed a consistent gap. People tend to understand what real-time APIs do pretty quickly, but seeing just how much you can build with real-time data takes a little more imagination.
When I demo Realtime Media Streams (RTMS) to someone for the first time, I can usually tell where their imagination is going. A stream of transcript data appears in the terminal, and the first ideas are usually familiar ones: summarizing the meeting, capturing action items, or using the conversation later with an AI assistant. Sometimes they wonder why they need RTMS at all if they can already record a meeting and process it afterward.
Using conversation data after the meeting is a natural place to start. In fact, I'd probably be thinking about the same things if I were seeing RTMS for the first time too.
People aren't misunderstanding the value of the data. They're imagining applications they've already seen. Meeting summaries and note takers are useful examples of conversational AI, but they're only one way to build with real-time meeting context. The more interesting question is what becomes possible when applications can use that context while the conversation is still happening.
A more useful way to think about RTMS is to look at how the pieces fit together: the conversation, RTMS as the pipeline to your application, the intelligence your application adds, and the experience that ultimately reaches the user. When that mental model is clear, the individual use cases start to look like variations on the same foundation.
Before information becomes structured
Before information becomes structured, it almost always begins as an unstructured conversation. A sales opportunity exists before it's entered into Salesforce. A treatment plan exists before it's documented in an electronic health record. By the time software enters the picture, people have already exchanged ideas, challenged assumptions, changed their minds, and decided what matters.
That has been the starting point for enterprise software for a long time because there wasn't another option.
Software has always been excellent at working with structured data. Conversations are different. They wander. People interrupt each other. Someone asks a question nobody expected. A meeting ends somewhere completely different from where it began.
Conventional software couldn't really participate in that process, so people had to translate the conversation into something software could work with. The conversation happened first. The structured record came later.
AI changed what people expect from meetings
AI has changed what people expect from meetings. It's become normal to leave a meeting and find a summary waiting for you, along with action items and follow-ups that didn't exist a few seconds earlier. As the models improved, those experiences became better too.
For many of the developers I work with, meeting summaries and other familiar AI experiences are the reference point when they first encounter RTMS. That's a useful place to start, but it can also make it harder to imagine applications they haven't seen before. If your mental model of RTMS becomes "the thing that powers meeting summaries," you've understood one application of the platform, but you've missed almost everything else it makes possible.
I still think of Realtime Media Streams as a good API to demo. But showing what it lets you build is much more fun.
Where the application takes over
RTMS has a very clear role in the architecture of conversational AI apps. It gives your application access to the conversation while deliberately leaving much of what happens next open to you.
This diagram became the simplest way I'd found to explain that division of responsibility.
I like this diagram because it makes the division of ownership clear.
RTMS opens up the conversation to your application in real time, as it unfolds. From there, you decide what happens next: which models to use, what context matters, what the application should understand, and where the resulting intelligence should go.
That's important because every organization interprets the same conversation differently. A sales team isn't listening for the same things as a healthcare application, and even two companies in the same industry may care about different signals, workflows, and internal knowledge. RTMS gives you access to the conversation, but the intelligence built on top of it is yours. You own the AI, the context, the business logic, and the decisions your application makes with the data.
Connecting the dots
At this point, it's probably easier to show you than to keep describing it.
Arlo is an application built to connect the dots. We needed something that demonstrated what building conversational AI apps looked like on Zoom.
Arlo uses the same real-time architecture to power five different experiences inside the meeting, showing how one foundation can support completely different kinds of applications.
That same idea became the thesis of my 2026 Developer Summit session : RTMS provides the stream. Developers create the intelligence.
There's one more piece of the architecture worth looking at: how the intelligence gets back to the user in the moment it matters. In Arlo, that happens through Zoom Surface Apps, which allow web apps and workflows to appear directly inside the Zoom meeting alongside the conversation itself. Surface Apps aren't required to use RTMS, but for Arlo they became a natural way to close the loop by bringing the application's intelligence back into the conversation.
Watch a 3-minute walkthrough of Arlo in action
Getting real-time conversation data into an application can now take just five lines of code . That means developers can spend less time collecting information and more time deciding what to do with it.
Build what matters to you
The questions changed, too. Instead of asking what RTMS could do, developers started asking what they could build.
There isn't a single answer to that question, and there shouldn't be. The most compelling applications I've seen started somewhere different because they were solving problems that already existed inside a particular business.
A useful place to start isn't the model or even the application. It's the conversations that happen inside the organization. Every business has conversations where people are making decisions, sharing context, and pulling information from different systems before they decide what to do next.
When helping people imagine what to build, I keep coming back to three questions:
- What conversations already happen inside the business?
- What decisions are people making during those conversations?
- What information would help them make those decisions before the conversation is over?
Those questions don't point toward one application. Every organization has its own knowledge, workflows, and ways of making decisions. RTMS doesn't prescribe what your application should become. It gives developers ownership and the opportunity to build around conversations that are already happening, using valuable context that only their organization has.
The questions we heard reflected that shift. Enterprise teams started describing problems they already wanted to solve. Sales teams wanted to stop switching between tools during customer conversations. I see teams looking differently at software and workflows they've used for years now that they have real-time access to their own conversation data. They can connect that data to the systems and context they already have, decide what matters for their business, and build around the way their business actually works.
At that point, the conversation isn't really about what the API does anymore. It's about the problems developers want to solve with it.
What will you build?
Conversations are becoming programmable. The question is no longer whether you can build intelligence into your meetings. It's what kind of intelligence matters most to your business.
Get started with RTMS and see what you can build.

