OpenAI expands its AWS partnership: managed agents, ChatGPT in Slack and Teams
OpenAI will run managed agents on AWS and bring ChatGPT to Slack and Teams without individual licenses, with a private-inference preview coming this fall.
Published entries across all sections carrying the “Model inference” tag, newest first by publication date on this site.
3 entries
OpenAI will run managed agents on AWS and bring ChatGPT to Slack and Teams without individual licenses, with a private-inference preview coming this fall.
Inferact, founded by the original vLLM team, wrote a megakernel inference kernel for Google TPUs. Paired with DeepSeek’s DSpark speculative decoding, 16 TPU v7 chips served Kimi K3 at 709 tokens per second versus 452 on GB200 under the same setup, QbitAI reports. The code is open source.
Free online chat access and WebGPU-accelerated in-browser local inference for the open-weights DeepSeek R1 reasoning model.