#multimodal
Loot, blog posts and adjacent themes connected to this topic. Follow the tag to keep it in your orbit.
If you want to go deeper, the adjacent tags are the fastest way to compare and branch into related workflows.
More from this topic
Related reads
Z.ai releases GLM-5.3-Flash with open weights for multimodal agents
Z.ai has released GLM-5.3-Flash with public weights, native multimodal inputs, a 1M-token context window, and a focus on coding and long-hor…
DeepSeek V4.1-Flash replaces V4 Pro routing with cheaper multimodal API access
DeepSeek V4.1-Flash is live with native image understanding, lower token rates, and a September 14 routing change that moves V4 Pro requests…
OpenAI launches ChatGPT Images 2.5 with two API-ready variants
OpenAI has announced ChatGPT Images 2.5 for more polished image creation and editing. Vercel independently lists two OpenAI variants for AI …
Google puts Lyria 3.5 into the Gemini API for full-song generation
Google's Lyria 3.5 music model is moving from a product showcase into the Gemini API public preview, with full-song generation, vocals, lyri…
DeepSeek adds vision to V4 Flash at experimental API launch
DeepSeek's experimental V4 Flash Vision model now accepts images through its API at the same listed token rates as V4 Flash, while independe…
MiniMax releases M3 with 1M context and native multimodal input
MiniMax has released M3, a 428B-parameter open-weight model that combines a 1M-token context window, native text-image-video input, and agen…
Qwen-Image-3.0 launches with long-prompt image generation
Alibaba's Qwen team has announced Qwen-Image-3.0, a new image model focused on dense layouts, small text, multilingual rendering, and design…
Gemma 4 12B Brings Local Multimodal Agent Workflows to Laptops
Google's Gemma 4 12B gives developers an open-weight, encoder-free multimodal model for local agent workflows on high-memory laptops, with L…
