
Fold-model trunks are memory-bound on triangle updates; this kernel folds roughly 1.4× longer sequences on the same GPU behind a one-line patch, and demonstrates CuTe DSL is enough to ship a production kernel without thousands of lines of CUDA.
reponanoAlphaZero – Train a grandmaster-level chess model in 24h with TPUstdoubleu
articleRealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI FrameworksJinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu
repounsloth v0.1.50-beta — Introducing AMD supportEtherllChecking sign-in…
Loading comments…