
Scoring one artifact with an is easy, but running rubric-based assessment reliably at scale requires infrastructure most teams build from scratch; OASIS packages that as a reusable CLI plus stack for anyone building pipelines.
“OASIS pairs a standalone command-line interface with a canonical integrated Elephant + MAPLES stack for encounter management and multimodal grading orchestration.”
“Distinctive features include rubric-as-program compilation, progressive execution plans, content-addressable grading identity, transcript-augmented multimodal grading, explicit review state, and a shared command surface for humans and autonomous agents.”
“In production at UT Southwestern Medical Center since Fall 2023, the platform has processed more than 7,000 encounters.”
articleDo large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effectsEmad Alharbi
articleolmo-eval: An evaluation workbench for the model development loopallenai.org
articleCompliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesJuan Yeo, Geewook KimChecking sign-in…
Loading comments…