Choosing a model using workload evidence and operating constraints.
Model selection balances task quality, latency, context needs, modalities, safety, availability, and cost. A global benchmark rank is weaker evidence than an evaluation set representing the actual workload.