OpenAI vs. Deepseek vs. Qwen: Comparing Open Source LLM Architectures
Source
youtube.com
Author
Y Combinator
Date
Why it matters
Shows how GPT OSS is built and how it differs from DeepSeek and Qwen, which helps when choosing an open weightsA model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.Full definition → model for inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → cost and deployment.
Terms in this piece · Glossary
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.