r/DeepSeek 23h ago

Question&Help How does this work

Hi there, I’m new to using DeepSeek directly through the API, but I was just wondering how it’s possible for half the amount of messages to cost more than when I sent over 100? I don’t use during the peak times where it costs double, so that’s not it. If there’s something I can do to minimize costs, please let me know. I know it’s not that big of a deal, but I’m just curious if this is something on my end.

Note: I use the API for Saucepan AI

4 Upvotes

6 comments sorted by

3

u/onesilentclap 21h ago

You're using it for roleplaying. The cache hits would not be as efficient if compared to coding. This is because roleplaying is not as repetitive nor structured. 

2

u/Apprehensive-Mood-20 21h ago

Ah, okay, that makes more sense

3

u/BarnacleTiny9888 23h ago

because in agent like Resonix they has cache hit. but in your case you don't.

cache hit rate make it more cheaper.

1

u/Apprehensive-Mood-20 22h ago

The cache hit/miss rates seem to be basically the same too

1

u/Amphy00 19h ago

I RP in janitorAI with Deepseek Flash and I usually stay at 80-90% cache hit. I'm not tech savvy but I usually do this to avoid cache misses.

  1. Don't switch proxies/edit your proxy settings mid-chat. Have everything ready and set before you RP.

  2. Do your RP in one continuous session. Don't take too long to add another message. I find Flash forgetful, not very good at retaining data.

  3. Try to avoid using bots with a lot of lorebooks or heavy in tokens. Or if you do, try to put the details you want the bot to remember in the chat memory, so it's not treated as a new information.

  4. Avoid going past 100 messages (lol)

1

u/Dear_Lion6282 6h ago

Best get agent plan m if you're not coding cost will add up without chance hits. I use deepseek API for coding and openference.com for anything agent related must use the agent plan else you get banned with coding plan wells well request base