
A rare prototype-to-deployment account of a knowledge measured against a strong baseline with 75 real professionals, including which weaknesses forced rework: multi-document retrieval, answer validation, and targeted proactive dissemination.
articleResearch Assistant: AstraZeneca's Agentic System for R&DPiotr Grabowski, Mohamed Alameen, Jorge Bretones, Sabina Cardell, Miguel Carmona, Gavin Edwards, Ben Grainger, Sameh Hassan, Erik Jansson, Artur Kuziakhmetov, Albert Maristany, Hebatallah Mohamed, Andriy Nikolov, Sebastian Nilsson, Mark O'Donoghue, James Pacileo, Ashiq Sultan, Alex Voegele, Michael Ughetto
articleBC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERPHaoran Sun, Klaus Marius Hansen
articleAI Evaluation Should Work With HumansJan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond DouglasChecking sign-in…
Loading comments…