# Highlights - September 2026

By
Sebastien Frenck
●
Published 2026-09-01
  • Added inference-glm53-flash (GLM 5.3 Flash) to Active Models.

  • Added inference-glm5 (GLM 5.3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: zai-org/GLM-5.3.

  • Added inference-kimi-k3 (Kimi K3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: moonshotai/Kimi-K3.

  • Added inference-minimax-h3 (MiniMax H3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: MiniMaxAI/MiniMax-H3.

  • Added inference-qwen3-embedding-8b (Qwen3 Embedding 8B) to Active Models on the Phoeniqs Model Service. Hugging Face reference: Qwen/Qwen3-Embedding-8B. Embedding model optimized for text retrieval across 100+ languages.

  • Added inference-qwen3-reranker-8b (Qwen3 Reranker 8B) to Active Models on the Phoeniqs Model Service. Hugging Face reference: Qwen/Qwen3-Reranker-8B. Reranker model for relevance scoring across 100+ languages.

  • Added inference-flix-swissgerman-full (Flix Swiss German Full) to Active Models on the Phoeniqs Model Service. Hugging Face reference: Flix-AI/flix-swissgerman-full. Swiss German ASR model that transcribes dialect speech into Standard German text; pricing matches WhisperX (inference-whisper-large-v3).

  • Launched PHOENIQS Data Pools, the retrieval layer for PHOENIQS Chat and AI agents; see the new PHOENIQS Data Pools page. Hybrid RAG (vector plus full-text) over unstructured documents and CSV/TSV data, PostgreSQL Row-Level Security scoped to the user's identity, and read-only MCP and REST access. Default per-user storage quota of 5 GB, configurable per deployment.

  • Launched the Chat Dedicated plan (CHF 260/month excl. VAT, 10 GB RAG storage, single-tenant instance), replacing the Chat Business plan on the PHOENIQS Chat Subscriptions page. Chat Dedicated does not bundle AI credits; model usage is billed through your Base LLM API subscription. Existing plans remain unaffected.

  • Added Provision PHOENIQS Chat, a step-by-step guide to provisioning a dedicated Chat instance from the Phoeniqs Console.

  • Scheduled inference-glm45-air-110b (GLM 4.5 Air 110B) for decommissioning on 26.10.2026; see Scheduled for Retirement. Suggested replacement: inference-glm5 (GLM 5).

  • Scheduled inference-gemma-12b-it (Gemma 3 12B IT) for decommissioning on 26.10.2026; see Scheduled for Retirement. Suggested replacement: inference-gemma4-31b (Gemma 4 31B).

  • Decommissioned inference-apertus-70b (Apertus 70B) on 21.08.2026; moved to the Decommissioned Models list. Suggested replacement: inference-apertus-v15-70b (Apertus v1.5 70B) on Active Models.

  • Added two IaaS FAQs: storage encryption at rest (enforced by IBM Storage FlashSystem 9000, transparent on-the-fly encryption) and PVC snapshots/backups encryption at rest (with Mount10 backups).

  • Added one Key Management FAQ on Phoeniqs Cloud KMS availability: coming soon and not yet available, versus Cloud HSM, which is available in single or high-availability (HA) setups.

  • Added Programmatic API Access, a guide to authenticating and interacting with the OpenShift API using tokens and CLI tools.

  • Added Tenant Image Registry with Quay, a guide to deploying your own Quay registry in a namespace to build, store, and pull container images.