

They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.





Impossible to know.
There are systems for doing things cryptographically secure, but I don’t know much about that.