Author: Long Yue, Wallstreetcn
Three years ago, Meta CEO Zuckerberg said on the Joe Rogan show that one day when people put on glasses, an AI Agent would appear with them. Now, this scene may be happening.
Meta's personal AI Agent, Muse, reached millions of users within two weeks of launch. Zuckerberg said in an interview program on September 25: "Every few years we get something like this." He characterized Muse's early reception as "a home run right out of the gate"—which is rare in Meta's product history.
At the same time, he announced that Muse will be integrated into the entire line of Ray-Ban smart glasses, and users will no longer need to say "Hey Meta," but can customize a wake word to directly call their own personal AI Agent.
For years, the metaverse, smart glasses, and large models, which the outside world viewed as three independent gambles, are accelerating their convergence at Meta's internal product level.
In this interview, Zuckerberg also admitted that the failure of Llama 4 was the "scariest moment," and revealed that Meta is building a 5-gigawatt training cluster, believing that sufficiently large compute can "brute-force" AGI.
Why Muse can be "a home run right out of the gate"
Muse's starting point was Zuckerberg sitting down with team members Nat and Alex, using open-source tools to "piece together" an early version at home.
"We realized this was a magical experience," Zuckerberg said, "if we could make it something anyone can use—without having to buy a Mac Mini yourself, without having to tinker in the terminal—then this would be something billions of people would want to use."
This judgment drove the entire subsequent R&D logic: not just training models, but full-stack self-development from the model, the scaffolding, to the Agent's "heartbeat" mechanism—the Agent will automatically wake up periodically, check the user's goals, and proactively advance to-do items.
Post-launch data confirmed this judgment. Two weeks, millions of users.
Personal AI: the focus is execution, memory, and "judgment"
Regarding the difference between personal AI and general AI, Zuckerberg summarized the current focus of competition as making models better Agents.
"The biggest thing in the past year is basically coding Agents." He said coding ability is important because even if users are not directly writing code, "your Muse will also be writing code for you in the background all the time, to accomplish various things."
But he believes that a personal Agent cannot only have coding or task execution capabilities; it also needs to understand privacy boundaries and social context.
Zuckerberg gave an example: when a user asks Muse to book a restaurant, the Agent may know information such as the user's allergies or pregnancy, but that does not mean this information should be automatically disclosed. "You want it to complete the task while disclosing as little information as possible."
He said this capability belongs to the "basic social skills or common sense" that humans usually possess, but if companies are only training coding Agents, many of them have not incorporated it into the model's capability scope.
"We train the entire model, rather than taking someone else's ready-made model and building a framework and Agent around it," Zuckerberg said.Meta has also added features at the product level such as memory, task tracking, virtual character animation, and real-time voice interaction.
On personalization, he said Meta does not believe AI should have only one fixed personality. "Many labs are focused on how to tune personality to the right state, but I have never believed there is only one correct answer for personality."
He said Muse allows users to modify the avatar, voice, and basic personality, and the model is designed to be highly controllable. "You can define how you want to interact with it, which is very important for a personal Agent."
Every Agent has its own "computer"
One of the core infrastructure differences between Muse and other AI products is that Meta equips each Agent with its own Secure VM.
Zuckerberg explained why this is necessary:
Your Agent will know a lot of sensitive information about you. We do not believe this information should be mixed into one pool with everyone else's data.
He compared the Muse Secure VM to "that computer under your desk that belongs only to you"—user data is stored encrypted, and sensitive credentials such as passwords are managed through an independent security module that the Agent itself cannot directly read.
On this basis, Meta has also designed a "Sentinel" security Agent specifically to monitor data flows in and out of Muse. Once anomalies or high-risk operations are detected, it will directly intercept them and prompt the user for authorization.
The metaverse, smart glasses, and Agents: three paths are converging
In the interview, Zuckerberg revealed that Muse will be fully integrated into the Ray-Ban series of smart glasses and will upgrade the existing interaction method.
The current glasses offer a "single-turn conversation" experience—say one thing, get one reply, and it ends. After integrating Muse, the glasses become the front-end entry point for the Agent: the user speaks, and the back-end Muse continues working in the Secure VM, then reports back once the task is complete.
Regarding the product roadmap for the glasses, he described a clear hardware ladder:
- Camera-free audio-only glasses: already equipped with Muse, can handle calls, music, and voice tasks
- Meta Ray-Ban with a small display: already released, with basic visual feedback
- Full-field-of-view holographic AR prototype: already released, which Zuckerberg called "very exciting"
He said Meta's long-term investment in glasses puts the company in a relatively favorable position once AI Agents mature.Meta is bringing Muse and more AI features into its glasses products, while metaverse technology, which previously emphasized a greater sense of "presence," is still advancing, though more resources are currently shifting toward Muse and the AI features of smart glasses.
As for virtual avatars, Zuckerberg said Meta previously needed room-scale devices, multi-angle scanning, and enterprise-grade GPU environments to generate relatively high-quality realistic avatars; now, users can complete setup with just a few photos, and the related capability can already run in a pair of VR glasses.
Looking ahead to 2030, Zuckerberg said that the core vision of the metaverse has always been to integrate the physical world and the digital world. For example, he said, in the future people can participate in activities offline together with friends who join via holographic images; in work scenarios, humans and multiple Agents can also participate together in group chats or meetings, and Agents can appear as holographic figures or in other embodied forms.
"The Scariest Moment": Llama 4's Misstep
Not all bets have gone smoothly. In the interview, Zuckerberg rarely spoke directly about the failure of Llama 4.
After Llama 4, that was the scariest moment... I thought we were on track, but we weren't. It was a fairly big negative surprise.
He attributed the problem to a fundamental mistake in team structure:
I built the team in the way of an Instagram recommendation system or an ads system - hundreds or thousands of people pushing forward in parallel. But training a language model requires a tightly collaborative small team, doing it as a group science project. Every seat is extremely valuable.
As a result, Meta completely reorganized, broadly bringing in top talent from across the industry, and established the Meta Super Intelligence Lab (MSL). Zuckerberg said that a new generation of models will be released soon, but it will not be announced at the Connect conference.
Computing Route: Brute-Forcing AGI, 5-Gigawatt Project Under Construction
When talking about the technical path to AGI, Zuckerberg gave a direct judgment.
I'm not sure what fundamental architectural breakthrough is still needed... I think we already roughly know the recipe. If you can build a large enough supercomputer cluster, you can brute-force your way there.
Meta's current computing expansion route: a cluster of more than 1 gigawatt in Ohio is basically already in use for training the next generation of models; a 5-gigawatt-class cluster in Louisiana is under construction.
He said, "When you have multi-gigawatt clusters for training, you will basically get something close to AGI or even superintelligence."
However, he also added that the current computing-driven route does not mean that architectural research is unimportant -
The human brain consumes only about 10 watts, while our systems may be a million times less efficient than it. Only by combining large-scale computing power and architectural breakthroughs can we truly lead.
Alignment Is Not a Burden, but a Problem the Product Must Solve
Zuckerberg's attitude toward AI safety is very pragmatic - he does not regard it as regulatory pressure, but as a prerequisite for whether the product can succeed.
If you ask Muse to do one thing, but it does the opposite, who would still use it? We need to make the model not only understand your specific instructions, but also understand your intent and your values.
"For Muse to reach a billion users, we must solve the alignment problem, or at least make very great progress on it." he said.
Regarding the safety boundaries during the training process, he used an analogy: "It's like parents setting rules for their children—if it 'completes' a programming problem by modifying system configurations, I want to tell it: No, I asked you to learn the problem-solving method, not to find a shortcut to bypass it."
The full interview is as follows:
Making Big Bets
Host: Billions of people will want to use this. For you, what does it feel like to be building it? Is 'immortality' a possibility? Okay, if you take me into your mind and fast-forward to 2030, what would that look like?
This is Mark Zuckerberg. Twenty-two years ago, he built a social network that connected billions of people and forever changed the world. And now, he has decided to build something even greater. To that end, he is making big bets in the metaverse, smart glasses, and artificial intelligence. For years, skeptics thought these were three separate, impossible bets, but they missed the bigger picture—because right now, these bets are converging to jointly create a brand-new kind of superintelligence.
Today, we will present all of this to everyone, and I will also ask Mark some questions he has never been asked before, listen to his vision for the future, so that you can plan ahead and build the next big thing.
Host: Thank you very much for coming on the show.
Mark: Thank you, I'm happy to be here.
Host: For this conversation, I watched every interview you've ever done.
Mark: Wow, that's more than I've watched myself.
Host: It was fun, and very rewarding. Two things left a deep impression on me: first, your love for 'building'—I feel you are one of the top builders; second, your ability to make bets—the courage to make huge bets. I think this week, all these bets are converging, so let's start there.
Mark: Okay. We've been doing these things for a long time. On AI, as a company, we've been doing it almost since the beginning—the first version of the news feed was, in a way, a machine learning product. Then we created the AI research lab about 15 years ago.
But now we've entered a new phase—about a year or so ago, we launched Meta Superintelligence Labs. It was a fairly thorough research restart, bringing in a lot of great talent from across the industry, which is exciting. Currently, we're seeing models get better and better, and the next generation of models is about to be released, though it won't be announced at Connect. In addition, we have the Muse personal assistant, which so far has been very well received.
When building these things, you actually aren't sure how it will turn out. We ourselves like it. Earlier this year, I put together Open Claw at home in my own way, to feel how this thing works and how to turn it into a magical experience that anyone can use.
Basically, when we—me, Nat, and Alex—sat together and realized this was a magical experience, and that if we could turn it into an out-of-the-box version that ordinary people who don't understand technology, don't want to install a Mac Mini themselves, don't want to mess around in the terminal, and don't want to debug when something goes wrong could also use, I felt that this would be something billions of people would want to use.
Since then, we've been working toward this goal: tuning models specifically for it, building not only the assistant itself and the runtime framework, but also the technology that provides each assistant with an independent computing environment—we built the entire Muse secure virtual machine for this.
Internally, we always thought this was special, and the internal team loved it, but you never know how people will react after a product launch. Occasionally there are cases where it hits a home run right out of the gate, but most of the time you get some positive feedback and need to iterate on a few things before the product really "clicks." And this time, it "clicked" right away. It's really exciting to see that happen—in just two weeks, millions of people are already using it, which is quite rare. We only encounter this kind of situation every few years, but it's definitely one of the most enjoyable moments for a startup.
Host: I really admire your ability to stay in the game and keep trying. And hitting a home run so many times is truly impressive. In a previous interview with Joe Rogan, about three years ago, you mentioned at the end that one day you could just put on glasses and an AI assistant would appear. So I feel that Muse already has a physical form, which is very smart because it feels inevitable.
Mark: I think it also just makes it seem friendlier and cuter. I think too many people describe AI as something terrifying, but AI should just be useful and fun. That physical form—early in the project, a designer created that character, and for some reason they kept wanting to iterate, but the first version was the best. Later someone said, "Oh, it has to be blue because it's Meta." I said, "No, I think this form is right, you nailed it the first time." And that's it, it's just fun.
How is personal AI different from general AI?
Host: Great detail. If you want the model to excel in personal matters, not just as a general intelligence model, how would the training approach differ?
Mark: I think the core right now is making the model an excellent "agent." Over the past year, the biggest trend has been coding agents, which contain two core ideas: coding expertise and the general ability to be a great agent.
Our strategy is to prioritize making the agent itself great, rather than specializing in coding ability.
Coding ability is important because even if users don't think they're writing code, Muse is always writing code in the background to complete various tasks. But we believe the agent should first and foremost be an excellent agent, and coding ability serves that goal.
Moreover, when building a personal assistant, some things are more important than building enterprise software products. For example, if you're building an enterprise coding tool, the model doesn't need to have any concept of "information disclosure boundaries"—like what information should be said and what shouldn't. But for Muse, this is very important.
You need to train this ability into the model, just like any other ability. You tell it many things, and then you want it to help you achieve your goals. For example, you want it to help you book a restaurant, you're looking for a suitable restaurant, but you might have some allergy, or you're pregnant, and you might not want to disclose these to the restaurant. But Muse will know this information, and you want it to complete the task while revealing as little information as possible. This is a specific skill, essentially basic human social common sense.
But most companies that only do coding agents haven't trained this ability into their models. We can do this because we're doing full-stack—we're not taking an existing model from someone else and putting a framework on it; we're training the entire model from scratch, specifically for these abilities.
The model certainly needs broad general intelligence, but it also has these specific abilities. Then on top of that, you build the entire assistant and all the details around it: memory, the runtime framework, and the way it "heartbeats"—it periodically wakes up and checks, "Okay, these are the goals I know about you, is there anything I can advance right now?"
We also have a dedicated team solely focused on polishing the real-time animation effects of the avatar, because it's not just the default avatar—you can customize any avatar, and then it moves naturally, and the effect is fantastic. We're also launching voice mode, where you can have real-time voice conversations, and your assistant is right there with you. These details, I think, all stem from us doing full-stack—the model and the product are developed together.
Host: Another big thing is the virtual machine. Can you explain why it's important and what it unlocks?
Why does Meta give each AI assistant its own computer?
Mark: Basically, for an agent to do things for you, it needs a place to store your information. We believe this information shouldn't be mixed with everyone else's information in a shared resource pool.
Think of it this way: your assistant will know a lot of sensitive information about you. Many people initially encounter agents like OpenCloud by setting up a Mac Mini at home. So we thought, many people don't want to buy a Mac Mini or set it up themselves. So what kind of experience would best approximate that effect? The answer is: you just download an app, sign up, and get a computer dedicated to you, for your assistant to work and store data, and build a security model around it. This approximates having your own computer under your desk—even we at Meta can't access it.
For example, Muse confidential virtual machines—this is a feature we are developing, and even we ourselves cannot see the contents inside your virtual machine.
We have also built a lot of things around these, such as secure credential storage—when your assistant needs to handle information like passwords, it itself does not need to be able to see these passwords; it only needs to be able to "insert" the credentials when you ask it to log in to a certain service, and only do so when you explicitly ask. The system should be designed this way: that information itself should not be freely accessible, because accidents always happen, someone may try to break in, or the system may have problems.
So you need to ensure that the agent cannot access this data, and Meta cannot access it either. Providing each assistant with an independent computer and making security as extreme as possible is the fundamental foundation of this technology—it both gives Muse the capabilities needed to help users achieve their goals, and ensures privacy and security, making it a world-class, industry-leading product in this field.
Host: So does this mean that, just like WhatsApp, the data on the virtual machine that Muse connects to is encrypted? How should people understand how data is actually stored in the virtual machine?
Mark: We basically built two versions. The Muse Secure VM includes various privacy features, including the entire Sentinel agent architecture we built. You have an ordinary Muse assistant that is executing tasks for you, and at the same time there is also a security agent we call Sentinel, specifically monitoring the data flow in and out of your Muse.
If external content tries to break security, Sentinel will directly cut it off and stop it from happening. If it believes your Muse is about to take an action that requires your intervention, it will override Muse's operation and trigger an alert for a person to confirm permission—for example, "Do you want Muse to be able to do this?"
This entire system, plus secure credential storage, plus multiple layers of defense in depth, together make up the Muse Secure VM.
We are also developing another project. Nat and I specifically recruited Moxie Marlinspike—the very person who worked with us back then to implement WhatsApp end-to-end encryption—to design the Muse Confidential VM. The core idea is that, on top of the secure virtual machine, you are also given a dedicated encryption key, so that even Meta cannot access the contents inside.
This version is more difficult to implement, because if Meta cannot access the inside of the virtual machine, debugging and ensuring the system runs properly becomes much harder. So we spent some time on it, but it will be launched soon. This will basically reach the security standard people are already familiar with on WhatsApp and our other most secure products.
Host: Is the advantage of doing this just making people feel psychologically safer, or are there actual benefits?
Mark: I think security itself is important. Our goal is to approximate giving you a local machine under your desk. What does that local machine give you? It means no company can access it.
So, assuming Meta wants to provide you with this service, how can we give you the same level of privacy and security guarantees, so that no company—whether Meta, or anyone trying to break into us, or in some countries where you do not trust the local government—can access it? Because we ourselves cannot get in either, because we do not have access.
I think this is very important, and it is also an important reason people trust WhatsApp. This is real value for privacy, security, and trust.
If you are going to have an assistant that knows everything about you—I guess almost all of us will have such an assistant—fast forward 5 years, everyone will have an assistant that deeply understands your goals and everything, and can help you get things done. In this situation, being industry-leading in privacy and security is very important. We wanted to do this from the very beginning.
How to shape the personality of AI?
Host: There is another thing that is very interesting—you studied psychology in university.
Mark: Well, I was only there for a very short time, two years, but I feel it influenced many things I later built.
Host: When we look at models, we often say that a certain model is "very smart." But just as we choose friends, it is certainly because they are smart, and also because we like their energy and the way we get along. How do you think about shaping a model's personality?
Mark: I think the ideal model should be adaptable enough to fit different people's styles. I think many people in the industry have got this wrong—many other labs are focused on "how to design the personality right," but I never thought personality is a fixed thing. This is one of the reasons I so strongly believe in open source, so strongly believe that people can customize, and it's also why we designed Muse as a highly personal product.
You can customize and personalize Muse—not just the appearance and voice. When you first sign up, the first thing it asks you is "what do you want my basic personality to be like," and you can change it at any time.
We strive to make the model highly steerable, so you can define the kind of interaction you want. This is an important part of making it an excellent personal assistant—this adaptability around personality is crucial.
Host: What style is your own Muse?
Mark: I made it direct and efficiency-focused. It's quite interesting. Earlier versions were very sarcastic and humorous, but the current version is more straightforward. My assistant uses the default Muse appearance, but I dressed him in a toga and gave him a deep, somewhat comical voice. Interacting with him is fun.
Host: I think to have a sense of humor, you must be really smart. Many people don't realize that comedians are some of the smartest people in society—quick reactions, enough wit. You definitely have that trait. Watching all your interviews, you've always performed well.
Mark's Unexpected Predictions About AI and the Metaverse
Host: In your interview with Theo, you talked a lot about the next frontier of technology and where all this is ultimately heading. The AI field has gone through several "winters," where people thought there would be no breakthroughs. The metaverse has also gone through several such periods, where everyone thought it was an undeliverable bet. I tried the new holographic avatar feature yesterday, it's very cool. In the interview with Lex, it felt like it would take 11 hours to record your face, now it only takes 3 minutes. How did this advance to where it is today?
Mark: In the overall development of the metaverse, when we founded Reality Labs, we always believed that eventually there would be normal-looking ordinary glasses, and over time, they could both provide immersive presence and become excellent AI devices—because glasses are the only form that allows a device to see what you see, hear what you hear, communicate with you by your ear all day, and eventually display images.
But 10 to 15 years ago, I assumed we would first achieve holographic technology, then highly developed AI. However, the path of technological evolution is interesting—we actually got AI and personal superintelligence first, and then the technology to make holographic technology widespread and affordable enough. This is something I didn't anticipate, but I'm glad we're working on both directions.
Our significant investment in glasses has put us in a very favorable position when AI assistants are ready. Many of the announcements at Connect were precisely about bringing Muse and a lot of AI features into glasses, and I think users will really like it. This is a big deal.
Regarding presence, we are still advancing, but progress is relatively slower because most of our energy has shifted to building Muse and AI features for glasses. However, we have a long-standing project on real-time high-fidelity avatars.
As you said, three or four years ago, you needed an entire scanning room to record a person from all angles, and then you needed enterprise-grade GPUs to render, which was very cumbersome. The demo we did for the Lex podcast was that kind of setup. Now we've basically made it run in a pair of VR glasses—this is the first glasses form factor that can achieve such an amazing VR experience, rather than a big headset. Just a few photos can create your avatar, and the progress is astonishing.
Host: And it can also drive expressions based on voice. In the demo, it made me laugh, made me interact, and then understood how my face moves with the audio track. You mentioned in that podcast that some people who are usually more reserved in their expressions actually want richer expressions in the virtual world. How do you view people distinguishing between their "virtual self" and "real self"?
Mark: I think we are still very early in understanding the sociology and psychology of this. I think people's perception of themselves and the image they want to project are often somewhat different from who they actually are. Since the birth of social networks, people have been carefully selecting avatars. With Muse's avatars, we've seen a similar phenomenon. It's not so much "curating" yourself, but "curating" the "person" you want to communicate with.
I think when you provide people with the ability to express themselves, you want it to truly capture them, and at the same time it is itself a form of expression, not just a pure mirror reflection—it is both communication and expression. We hope to build something that balances both. This is always an iterative loop: see how people use it, then improve. After years of work, we are now truly at the starting line—for the first time, we can truly put quite high-quality realistic avatars into products, usable on phones and in VR. I'm very excited to see the results.
What will 2030 look like?
Host: Okay, take me into your thinking, fast-forward to 2030, if everything goes well, what will the holographic technology side look like? Holograms plus Muse, plus glasses, how do they all come together?
Mark: My understanding of the metaverse vision has always been about effectively merging the physical world with the digital world. The basic idea is: we have this wonderful physical world, and at the same time we also have a wonderful digital world—the massive amount of content accumulated on the internet over 20 to 30 years, which is breathtaking. But the way we access it is either by sitting at a desk or through a small screen in our pocket, which is fundamentally very limited.
I think the most ideal version of this is a seamless fusion of the physical and digital worlds. Think of it this way: right now the two of us are here. In some future version, one of us might be a holographic projection, but you still feel that sense of each other truly being present, which is completely different from a video call.
And the core of virtual reality is precisely delivering this sense of presence—making you truly feel that you are in the same room with others, or in another place. You can achieve this with holographic technology, and mix it in various ways. For example, I can play a game of poker with friends, some of whom are present, and others who join through holographic avatars, and they can also play cards, and the card table itself can also be holographic, so that those who are not present can also be integrated into it.
At the same time, AI can also be given a physical form and appear in this scene. In work scenarios this makes a lot of sense—I am already using various coding agents to build things all the time. Imagine this: you have a group chat channel with several people and several agents, and you assign tasks to the agents. But sometimes everyone gathers for a meeting, and the agents should perhaps also be present. How do they appear? Simple, just add a few more seats on the couch, and they appear as holographic avatars. Or use that cute little Muse character, or a dragon, or any strange and weird image you create.
I guess this will feel quite natural in the future.
Host: Interesting. In your previous interviews you said that the tech industry often forgets about "fun." I think having a physical assistant appear there would also make it feel more real—as if work is really being outsourced. When you see Muse typing, you feel that something is really happening. This is done very well—being able to see what is happening in the browser. So do you think such a scenario is possible: you wear the glasses, and then control your computer, letting Muse do things for you on the computer?
Mark: Oh, definitely. VR can already do this. You can just find a place to sit down, a cafe is fine, and then open your workstation, with six monitors, write code on them, everything is there.
On the glasses side, the most popular model currently does not have a display, which on one hand makes it more affordable and allows more people to use it, and on the other hand we are still working to fit a display into the most compact form. But we have already released a display version of Meta Ray-Ban, which is very popular, and it is a small display. We have also released a prototype version of full-field-of-view holographic AR, which I think will be very exciting.
So the whole product line is like this: from pure audio glasses—no camera, looking just like ordinary glasses, but with Muse inside, allowing you to use various audio tools, listen to music, make calls—to higher-end versions, all of them exist.
Host: I am wondering, with those audio glasses, can you talk to Muse while having your home computer do things?
Mark: Yes, absolutely. We just released this feature at the Connect conference.Now all glasses are connected to Meta AI, which is a "single-turn" experience—you send a prompt, it replies, and that's it. But with Muse, we are basically going to upgrade all glasses to Muse. First, you no longer need to say "Hey Meta," you can give it any name, which is part of the fun itself. Then you just talk to it directly, and it connects to your Muse, and your Muse handles tasks in your secure virtual machine, helping you get things done.
Host: That's amazing!
Founder Mindset
Host: Okay, we are here now, Muse is progressing smoothly, and the glasses are also doing well, but about a year ago, many people were asking "what exactly happened with the superintelligence lab." At that moment, how did you feel inside? When things are not going well, but you still see the long-term vision, what kind of experience is that?
Mark: What really went wrong was the Llama project and Llama 4.
Llama 1 was a pretty interesting model, it started the whole open source AI movement, and we're very proud of that. Llama 2 achieved scale, Llama 3 was a good model, almost at the frontier at the time. Then came Llama 4, and we basically deviated from the trajectory we should have been on.
Whenever things don't go the way I expect, I spend a lot of time thinking: why did this happen? What do we need to change to do better? This time, my reflection was: I got the entire team structure wrong.
I modeled it the way we do machine learning work like Instagram feed or ads systems—with hundreds or even thousands of people working in parallel on many things. But for building language models, what you really need is an extremely tight small team, treating it as a group science project. You don't need many people, but that means every seat on the team is extremely precious.
So we gathered the best people from all over Meta, and at the same time brought in many brilliant people from across the industry, forming a brand new team—Meta Superintelligence Labs.
From my perspective, when MSL launched, I knew it would take some time to restart, rebuild the infrastructure, and train the next generation of models. But I knew we had assembled an excellent team, and if the team could gel and operate well, the results would be good.
For me, the most thrilling moment, was actually after the Llama 4 release—I thought we were on track, but it turned out we weren't. That was a fairly big negative surprise. I think, as an entrepreneur, you're always tested at these moments, because inevitably, not everything will go smoothly. And what truly determines the trajectory is: when things don't go the way you hope, how do you find a way forward.
Host: I guess the reverse is also true—when something far exceeds your expectations, like the launch of Muse, how do you make sure you can seize that opportunity?
Mark: Exactly. Now the whole company is all in—at first it was just a small team building a product, but now everyone realizes that this is really ready for the big stage. The whole company is thinking about how to make this thing big, how to let hundreds of millions of people experience it.
From optimizing all the infrastructure to make everything run smoothly and squeezing every bit of compute out of the existing GPUs, to various product teams embedding Muse in different ways—like glasses. Seeing everyone working together to make sure Muse can scale smoothly, that feeling is great.
How Mark writes and communicates vision
Host: As a founder, I feel that's the moment you yearn for most—everything coming together. How do you communicate the vision to the company? With so many things being pushed forward at the same time, I feel like you write a lot. What's your process?
Mark: Writing helps me a lot, both in refining my own thoughts and in communicating externally. This summer, I wrote a very long piece called "The Future Belongs to Everyone," about 15 pages, which helped me systematically








