OutOfHere 2 hours ago The biggest continuing limitation I see is that the cached input has to be at least 1024 tokens. This is terrible. It means a lot of good prefixes that are smaller will go uncached for no good reason. The threshold should have been 128.
The biggest continuing limitation I see is that the cached input has to be at least 1024 tokens. This is terrible. It means a lot of good prefixes that are smaller will go uncached for no good reason. The threshold should have been 128.
Guide: https://developers.openai.com/api/docs/guides/prompt-caching