SambaNova adds prompt caching for MiniMax M3 on SambaCloud
SambaNova has enabled prompt caching for MiniMax M3 on SambaCloud without requiring any code modifications.
According to SambaNova, reusing cached prefixes across long-context agent workflows delivers a 35% to 88% faster time to first token (TTFT), up to 4.7x overall speed improvements, and 90% lower input costs at $0.06 per million cached tokens, as outlined in its [announcement](https://bit.ly/4emDQ71).