Vibeleaderboard
Sesame

ai lab

Sesame

Sesame matters because it treats speech as a continuous interaction problem rather than a sequence of polished audio clips. CSM exposed a context-conditioned generation approach, while TurnBench gives researchers a shared dataset and protocol for measuring when a conversational system should wait, yield, continue, or respond.3,8,9

Profile

Overview

Oculus veterans return to voice and hardware

Sesame is a conversational AI and hardware company founded in 2023 by Brendan Iribe, Ankit Kumar, and Ryan Brown. Iribe co-founded Oculus, Kumar founded the spatial-computing company Ubiquity6 and later led work on Discord's Clyde assistant, and Brown held engineering leadership roles at Oculus and Meta Reality Labs. The team combines speech research with the product and hardware experience needed for an all-day voice interface.1,2,6

Maya and Miles demonstrate voice presence

Sesame became publicly visible in February 2025 through Maya and Miles, two research-preview voices designed to pause, breathe, vary delivery, and respond to interruption. The company calls the target voice presence: the sense that a spoken system follows not only the text of a conversation but its timing, emotion, and social context. More than one million people tried the preview in its first weeks, according to figures reported by TechCrunch from investor Sequoia.3,6

CSM makes part of the stack inspectable

The Conversational Speech Model, or CSM, is the public research artifact behind that thesis. It uses a Llama-family backbone and a smaller audio decoder to generate discrete audio codes while conditioning on the text and audio history of a dialogue. Sesame released a one-billion-parameter base checkpoint under Apache 2.0, while noting that a separate fine-tuned variant powered the interactive demo. The release made part of the approach inspectable, but TechCrunch found that the base model could clone a voice quickly and relied primarily on usage requests rather than technical safeguards.3,4,5

From a viral preview to agents and evaluation

Sesame has since moved from a web demonstration toward a personal-agent product and an evaluation program. It raised a reported $250 million Series B in 2025, opened a broader iOS preview in May 2026, and says lightweight intelligent eyewear is planned for 2027. In August 2026 it released TurnBench with academic and industry partners: a multi-domain, manually annotated conversational dataset, evaluation protocol, and leaderboard for end-of-turn and interruption prediction.6,7,8,9

Notable contributions

  1. 01Conversational context as speech conditioningCSM conditions delivery on previous text and audio turns so the generated voice can respond to conversational context. The contribution is Sesame's released model and evaluation framing, not contextual speech synthesis as a whole.3,4
  2. 02Voice presence as an integrated product targetThe Maya and Miles preview made timing, interruption, disfluency, emotional delivery, and continuity part of one public conversational experience rather than presenting voice naturalness as the only objective.3,5
  3. 03An open benchmark for turn timingTurnBench releases a multi-domain conversational corpus, standardized end-of-turn and interruption tasks, and a public leaderboard so timing can be compared separately from transcription or voice quality.8,9