

Was wondering because its cuda only, and just like freetokens is game changer for lower memory setups.
I vaguely remember some kv cache stuff in llama that could make performance similar to freetoken, but forgot what it was.
Was there a pr open for this? ( though i cant use it, no cuda )








It does to lawful evil/neutral/good people :p