The Strata briefing describes a local inference engine and prompt processing measurements reported by its author. Those figures can help frame an experiment, but they are not guarantees for another machine, model or workload.
Measure your application's complete path: context preparation, data processing, queueing and generation. Engine improvements can disappear with larger documents or concurrent sessions. Record the configuration and input sizes so results remain comparable.
Include privacy in the same experiment. A local model can still receive excessive data, and surrounding programs may write prompts to files. Reduce content to what is necessary, restrict access and inspect retention in every component.
Use fictional samples to compare answers, latency and the effect of sanitization. Rota's API supplies a policy layer before inference; it does not install this engine on your machine or guarantee national residency for every execution offered by the platform.