Introducing preemptible compute: the same compute, half the price
Source
www.together.ai
Date
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Half-price GPU capacity for interruption-tolerant jobs (experiments, inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → bursts, batch work) changes the cost calculus for teams running non-critical AI workloads on Together's infrastructure.