Vibeleaderboard
← All Intel
Intel / article

ERNIE 4.5 Gets a Major Inference Speed Boost

Source
unknown
Date
Terms in this piece · Glossary
  • attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters

PLAS is a sparse update that speeds up long- on ERNIE 4.5 models, giving practitioners a concrete method for cutting inference cost on long-context workloads.

Source link unavailable.unknown
Recommended reads
Comments

Checking sign-in…

Loading comments…