Inkling: Thinking Machines Lab’s Open Multimodal AI Breakthrough
Generated: 2026-07-17 · API: Gemini 2.5 Flash · Modes: Summary
Inkling: Thinking Machines Lab’s Open Multimodal AI Breakthrough
Clip title: Inkling: Why Thinky’s Open Model May Change “Everything” Author / channel: Prompt Engineering URL: https://www.youtube.com/watch?v=IB53DUrnYgI
Summary
This video introduces Inkling, Thinking Machines Lab’s first open-weights multimodal AI model, highlighting its significant impact on the AI landscape. The release is notable because, unlike many proprietary frontier models from Western labs (e.g., GPT, Claude), Inkling’s weights are open-sourced under an Apache 2.0 license, allowing users to download and own a copy rather than just rent access via an API. Furthermore, Inkling was trained from scratch with its own unique architecture, making it a truly novel offering rather than a remix of existing models. The video positions Inkling as the biggest open-weights model release from a Western lab (excluding Nvidia), filling a critical void in the market, which has previously relied on Chinese labs for comparable open-source options.
Inkling boasts a substantial architecture, featuring 975 billion total parameters with 41 billion active parameters, built on a Mixture-of-Experts (MoE) design. It leverages 256 experts, with six active per token, and operates with a 1 million token context window. The model was trained on an expansive 45 trillion multimodal tokens, encompassing text, images, audio, and video, making it natively multimodal from its inception. While benchmarks show Inkling is not strictly state-of-the-art across all individual metrics, it demonstrates remarkable performance as a well-rounded generalist, particularly strong in agentic tasks, factual reasoning, and design capabilities. A smaller preview variant, Inkling-Small, is also available, with 276 billion parameters (12 billion active), designed for more accessible deployment on consumer hardware.
Thinking Machines Lab also offers the Tinker platform, an API for researchers and developers to fine-tune models. Inkling is readily available on Tinker, enabling users to customize the model with their own data without needing to manage GPU infrastructure. The video showcased Inkling’s impressive ability to handle complex tasks, including generating detailed HTML/CSS websites based on prompts and creating real-time interactive dashboards like an ISS orbital tracker. This demonstrates its advanced coding capabilities and tool-use integration, where it can perform interleaved web searches and display its thought processes. The model’s system prompt reveals a knowledge cutoff of April 2026, indicating its up-to-date training.
In conclusion, Inkling represents a monumental step for the open-source AI community, particularly from a Western scientific perspective. Its open-weights, multimodal nature, and robust architecture provide an unparalleled opportunity for customization and innovation. While the current iteration might have specific performance and pricing considerations on the Tinker platform, its open availability empowers a broader range of developers and researchers. As the first major open-weights release from Thinking Machines Lab, Inkling lays a strong foundation, and future iterations are anticipated to push the boundaries of multimodal AI even further.
Video Description & Links
Description
In this video, I break down Inkling, Thinking Machines’ first open-weight model release under Apache 2.0, trained from scratch as a multimodal decoder-only MoE. I cover the core architecture (near‑trillion parameters, 1M context window, 45T multimodal tokens, 256 experts with 41B active params), how it compares on benchmarks (not SOTA overall but a strong generalist with notable design performance), and why it matters as a major Western open model release. I also walk through the Tinker platform for fine-tuning/post-training, mention the upcoming Inkling Small (276B/12B active), show playground settings like reasoning effort, context limits, and web search/tool calling, demonstrate website generation and a coding example (ISS tracker), and review pricing and options to run weights yourself.
LINK: https://thinkingmachines.ai/news/introducing-inkling/
My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0
Let’s Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: engineerprompt@gmail.com Become Member: http://tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0
Inkling: Thinking Machines’ First Open-Weight Multimodal MoE Model (Architecture, Playground Demos & Pricing)
00:00 Inkling 00:26 Why This Matters 01:05 Architecture and Scale 01:55 Inkling Small 02:57 Multimodal Strengths 04:06 Reasoning Effort Controls 04:36 Playground Tour and Prompt 06:50 Tool Use and Speed 08:13 Website Design Patterns 09:21 Coding Demo ISS Tracker
Tags
prompt engineering, Prompt Engineer, LLMs, AI, artificial Intelligence, Llama, GPT-4, fine-tuning LLMs
URLs
- https://thinkingmachines.ai/news/introducing-inkling/
- https://engineerprompt.ai/
- https://prompt-s-site.thinkific.com/courses/rag
- https://tally.so/r/3y9bb0
- https://discord.com/invite/t4eYQRUcXB
- https://ko-fi.com/promptengineering
- https://www.patreon.com/PromptEngineering
- https://calendly.com/engineerprompt/consulting-call
- http://tinyurl.com/y5h28s6h
- https://bit.ly/localGPT