
Understanding that GPT defaults to code/math, Llama to narratives, DeepSeek to religious content, and Qwen to exam questions helps engineers anticipate model biases when prompts are underspecified—useful for prompt design, model selection, and construction.
“Despite the absence of explicit task or topic specification, LLMs generate diverse content; however, each model family exhibits distinct topical preferences . GPT-OSS favors programming and math, Llama leans literary, DeepSeek often produces religious content, and Qwen tends toward multiple-choice questions.”
Together AI
“We use topic-neutral, open-ended seed prompts such as “Actually,” “Let’s think step by step,” or even just punctuation like “.” We remove chat templates entirely—no system prompt, no roles—and use standard decoding.”
Together AI
“GPT-OSS overwhelmingly defaults to programming (27.1%) and mathematics (24.6%). More than half of a model family’s output concentrates in these two domains!”
Together AI
“We find that GPT-OSS frequently produces advanced or expert-level content (68.2%) , such as depth-first search, breadth-first search, or dynamic programming (see Figure 3). Llama and Qwen skew much more toward basic or intermediate material.”
Together AI
“Degenerate text turns out to be one of the clearest windows into safety and privacy risks, precisely because it reflects uncontrolled generation . These behaviors rarely appear in standard benchmarks, yet they are highly revealing.”
Together AI
articleCan LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMsXinyi Wang, Hong Jiao, Ming Li, Sydney Peters, Hanna Choi, Tianyi Zhou, Qingshu Xu
postLLMs Get Better One Task at a Time Now, Not All at OnceAndrew Ng
articleOpen challenges in LLM researchChip HuyenChecking sign-in…
Loading comments…