← All IntelIntel / article
ERNIE 4.5 Gets a Major Inference Speed Boost
- Source
- unknown
- Date
Terms in this piece · Glossary
- attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
PLAS is a sparse update that speeds up long- on ERNIE 4.5 models, giving practitioners a concrete method for cutting inference cost on long-context workloads.
Source link unavailable.unknown
Recommended reads
- articleFastDeploy 2.0: A Large-Scale Model Inference and Deployment Toolkit with Native Support for ERNIE 4.5unknown
- articleERNIE 5.1 Officially Released! Topping Multiple Leaderboards — A Model That Writes Better and Understands You Moreunknown
- articleERNIE-4.5-VL-28B-A3B-Thinking: A Breakthrough in Multimodal AIunknown
Comments
Checking sign-in…
Loading comments…