#
Highlights - September 2026
Added
inference-glm53-flash(GLM 5.3 Flash) to Active Models.Added
inference-glm5(GLM 5.3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: zai-org/GLM-5.3.Added
inference-kimi-k3(Kimi K3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: moonshotai/Kimi-K3.Added
inference-minimax-h3(MiniMax H3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: MiniMaxAI/MiniMax-H3.Added
inference-qwen3-embedding-8b(Qwen3 Embedding 8B) to Active Models on the Phoeniqs Model Service. Hugging Face reference: Qwen/Qwen3-Embedding-8B. Embedding model optimized for text retrieval across 100+ languages.Added
inference-qwen3-reranker-8b(Qwen3 Reranker 8B) to Active Models on the Phoeniqs Model Service. Hugging Face reference: Qwen/Qwen3-Reranker-8B. Reranker model for relevance scoring across 100+ languages.Added
inference-flix-swissgerman-full(Flix Swiss German Full) to Active Models on the Phoeniqs Model Service. Hugging Face reference: Flix-AI/flix-swissgerman-full. Swiss German ASR model that transcribes dialect speech into Standard German text; pricing matches WhisperX (inference-whisper-large-v3).Launched PHOENIQS Data Pools, the retrieval layer for PHOENIQS Chat and AI agents; see the new PHOENIQS Data Pools page. Hybrid RAG (vector plus full-text) over unstructured documents and CSV/TSV data, PostgreSQL Row-Level Security scoped to the user's identity, and read-only MCP and REST access. Default per-user storage quota of 5 GB, configurable per deployment.
Launched the Chat Dedicated plan (CHF 260/month excl. VAT, 10 GB RAG storage, single-tenant instance), replacing the Chat Business plan on the PHOENIQS Chat Subscriptions page. Chat Dedicated does not bundle AI credits; model usage is billed through your Base LLM API subscription. Existing plans remain unaffected.
Added Provision PHOENIQS Chat, a step-by-step guide to provisioning a dedicated Chat instance from the Phoeniqs Console.
Scheduled
inference-glm45-air-110b(GLM 4.5 Air 110B) for decommissioning on 26.10.2026; see Scheduled for Retirement. Suggested replacement:inference-glm5(GLM 5).Scheduled
inference-gemma-12b-it(Gemma 3 12B IT) for decommissioning on 26.10.2026; see Scheduled for Retirement. Suggested replacement:inference-gemma4-31b(Gemma 4 31B).Decommissioned
inference-apertus-70b(Apertus 70B) on 21.08.2026; moved to the Decommissioned Models list. Suggested replacement:inference-apertus-v15-70b(Apertus v1.5 70B) on Active Models.Added two IaaS FAQs: storage encryption at rest (enforced by IBM Storage FlashSystem 9000, transparent on-the-fly encryption) and PVC snapshots/backups encryption at rest (with Mount10 backups).
Added one Key Management FAQ on Phoeniqs Cloud KMS availability: coming soon and not yet available, versus Cloud HSM, which is available in single or high-availability (HA) setups.
Added Programmatic API Access, a guide to authenticating and interacting with the OpenShift API using tokens and CLI tools.
Added Tenant Image Registry with Quay, a guide to deploying your own Quay registry in a namespace to build, store, and pull container images.