feat: bill OpenAI cache_write_tokens at cache-creation price with zero clamp
Parse OpenAI's native cache_write_tokens (chat prompt_tokens_details / responses input_tokens_details), bill it at the cache-creation ratio, and clamp the uncached prompt remainder at zero since cached + cache-write can exceed prompt_tokens. Propagate the field through chat/responses/claude format conversions and tiered expression billing (cc variable).
This commit is contained in:
@@ -81,6 +81,9 @@ var defaultCacheRatio = map[string]float64{
|
||||
}
|
||||
|
||||
var defaultCreateCacheRatio = map[string]float64{
|
||||
"gpt-5.6-sol": 1.25,
|
||||
"gpt-5.6-terra": 1.25,
|
||||
"gpt-5.6-luna": 1.25,
|
||||
"claude-3-sonnet-20240229": 1.25,
|
||||
"claude-3-opus-20240229": 1.25,
|
||||
"claude-3-haiku-20240307": 1.25,
|
||||
|
||||
Reference in New Issue
Block a user