
Defects that behave the same on CPU and GPU slip past the usual differential oracle; comparing equivalent operations across seven DL libraries catches APIs that quietly disagree on boundary and non-finite inputs.
articleHarnessing LLMs for Document-Guided Fuzzing of Python LibrariesBin Duan, Tarek Mahmud, Meiru Che, Yan Yan, Naipeng Dong, Dan Dongseong Kim, Guowei Yang
articleToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling AgentsYiShan Zheng, Yuan Wu, Yi Chang
articleThe Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored TestsDouglas J. LeithChecking sign-in…
Loading comments…