Topic

#multimodal

Loot, blog posts and adjacent themes connected to this topic. Follow the tag to keep it in your orbit.

#multimodal
1Shown loot
8Shown articles
8Linked neighbor tags
Topic paths

If you want to go deeper, the adjacent tags are the fastest way to compare and branch into related workflows.

Loot

More from this topic

Explore all loot

Run Thinking Machines Inkling through Vercel AI Gateway

0
Use Vercel's AI Gateway model page to test Inkling as thinkingmachines/inkling without adding another provider integration. Inkling is Thinking Machines' open-weights multimodal MoE model, with text, image and audio inputs, controllable thinking effort, and long-context support. Treat it as an evaluation target for multimodal agent workflows, not as an automatic replacement for stronger closed frontier models. Vercel has added Thinking Machines' Inkling to AI Gateway under thinkingmachines/inkling, so teams already using AI Gateway can test it through the same playground, routing, budget, retry and usage-tracking layer they use for other models. Inkling is an open-weights multimodal MoE model from Thinking Machines with text, image and audio inputs, controllable thinking effort and long-context support. Use it for practical evaluation of multimodal agent workflows, document/audio reasoning and fine-tuning paths; compare latency, cost and output quality before moving production traffic.
Free
Review open
0
Blog

Related reads

Browse blog
AI & Automation

Z.ai releases GLM-5.3-Flash with open weights for multimodal agents

Z.ai has released GLM-5.3-Flash with public weights, native multimodal inputs, a 1M-token context window, and a focus on coding and long-hor

AI & Automation

DeepSeek V4.1-Flash replaces V4 Pro routing with cheaper multimodal API access

DeepSeek V4.1-Flash is live with native image understanding, lower token rates, and a September 14 routing change that moves V4 Pro requests

AI & Automation

OpenAI launches ChatGPT Images 2.5 with two API-ready variants

OpenAI has announced ChatGPT Images 2.5 for more polished image creation and editing. Vercel independently lists two OpenAI variants for AI

AI & Automation

Google puts Lyria 3.5 into the Gemini API for full-song generation

Google's Lyria 3.5 music model is moving from a product showcase into the Gemini API public preview, with full-song generation, vocals, lyri

AI & Automation

DeepSeek adds vision to V4 Flash at experimental API launch

DeepSeek's experimental V4 Flash Vision model now accepts images through its API at the same listed token rates as V4 Flash, while independe

AI & Automation

MiniMax releases M3 with 1M context and native multimodal input

MiniMax has released M3, a 428B-parameter open-weight model that combines a 1M-token context window, native text-image-video input, and agen

Kreativ & Medien

Qwen-Image-3.0 launches with long-prompt image generation

Alibaba's Qwen team has announced Qwen-Image-3.0, a new image model focused on dense layouts, small text, multilingual rendering, and design

AI & Automation

Gemma 4 12B Brings Local Multimodal Agent Workflows to Laptops

Google's Gemma 4 12B gives developers an open-weight, encoder-free multimodal model for local agent workflows on high-memory laptops, with L