AIHubMix Blog

Announcements, tutorials, model analysis, and product updates from AIHubMix.

GLM-5.3 Hands-on Guide: Always-on Thinking, Three Effort Levels, and the API Support Matrix

GLM-5.3 Hands-on Guide: Always-on Thinking, Three Effort Levels, and the API Support Matrix

An August 2026 guide to calling GLM-5.3: always-on thinking with three reasoning_effort levels, reasoning summaries, parallel tool calls, structured output, and automatic caching — with tested examples for the AIHubMix Chat, Responses, and Messages APIs.

7 min readTutorial
DeepSeek V4 Pro (0813): Thinking Passback & 3-API Matrix

DeepSeek V4 Pro (0813): Thinking Passback & 3-API Matrix

DeepSeek V4 Pro (0813) hands-on guide: thinking toggle and reasoning_effort levels, mandatory thinking-history passback, tools, caching, and a 3-API matrix.

15 min readTutorial
DeepSeek V4 Flash Was Degraded Today. Here’s Why Multi-Provider Failover Matters

DeepSeek V4 Flash Was Degraded Today. Here’s Why Multi-Provider Failover Matters

On August 4, 2026, DeepSeek’s official status page recorded two API degraded-performance incidents. The first incident lasted 1 hour and 18 minutes, from 02:02 to 03:20 UTC, and affected DeepSeek V4 Flash, V4 Pro, and Expert Mode. The second incident lasted 36 minutes, from 03:43 to 04:20 UTC, and affected the DeepSeek V4 Flash API. Both incidents have since been resolved. OpenCode also reported that DeepSeek Flash was experiencing capacity issues due to unprecedented demand. However, DeepSeek

4 min readOpinion
June 2026 Release Spotlight: ~20 New Models

June 2026 Release Spotlight: ~20 New Models

In June 2026 AIHubMix added ~20 models (glm-5.2, minimax-m3, qwen3.7-plus, kimi-k2.7-code, Kling video) plus LLM Router, Mapping & Fallback, CLI, backup domain.

3 min readChangelog
Global Acceleration: 75% Lower Latency, 99.99% Availability

Global Acceleration: 75% Lower Latency, 99.99% Availability

AIHubMix runs a self-built acceleration network: 75% lower latency, 60% less fluctuation, 99.99% availability, minute-level probes, automatic failover.

2 min readAnnouncement
OpenAI Compatible Interface Upgrade: Deep Support for Claude

OpenAI Compatible Interface Upgrade: Deep Support for Claude

AIHubMix upgraded its OpenAI-compatible API for Claude: interleaved thinking with no extra parameters, prompt caching, and Anthropic beta feature support.

8 min readNews
GPT-5.6 Is Live: Prompt Caching Billing Changes Explained

GPT-5.6 Is Live: Prompt Caching Billing Changes Explained

July 2026: GPT-5.6 gpt-5.6-sol / terra / luna on AIHubMix: 1.05M context, 1.25x cache writes, prompt_cache_key, explicit breakpoints, vs Claude caching.

8 min readOpinion
July 2026 Release Spotlight: ~30 New Models and Media APIs

July 2026 Release Spotlight: ~30 New Models and Media APIs

AIHubMix added ~30 models in July 2026, including claude-opus-5, GPT-5.6, kimi-k3 and qwen3.8-max-preview, plus media generation, 3D generation and MCP Server.

6 min readChangelog
Free AI Models on AIHubMix

Free AI Models on AIHubMix

Free AI Models: The Ultimate Guide to Building with Zero-Cost AI in 2026 on AIHubMix

13 min readAnnouncement
Kimi K3 Hands-On Guide: New Parameters & API Support Matrix

Kimi K3 Hands-On Guide: New Parameters & API Support Matrix

July 2026 Kimi K3 guide: reasoning_effort max, thinking history, dynamic tool loading, structured output, auto caching, partial prefix, and vision inputs.

8 min readTutorial
Claude Opus 4.7 New Parameters Guide

Claude Opus 4.7 New Parameters Guide

This article covers two key changes to reasoning control in Claude Opus 4.7, along with complete usage instructions for both the AIHubmix native API and the Chat unified interface. See also: Anthropic official announcement and model change log.

3 min readTutorial