Deepseek V4.1 Flash


DeepSeek V4.1 Flash is DeepSeek's open-weights, ultra-efficient vision-language Mixture-of-Experts (MoE) foundation model. Built on a novel Causal Encoder-Decoder (CED) architecture, it distributes 552B total parameters while activating only 8B parameters during prefill and 16B parameters during decode. Featuring native visual-text joint embedding, Compressed Sparse Attention, Engram memory, and a 4x KV-cache memory reduction, it effortlessly processes up to a 1M-token context window at high throughput. Engineered for extreme cost efficiency and frontier agentic coding, DeepSeek V4.1 Flash powers autonomous software engineering, multimodal document RAG, UI browser automation, and high-concurrency API integrations.

deepseek-v4.1-flash

deepseek-v4.1-flash is an efficiency-optimized Mixture-of-Experts (MoE) model developed by DeepSeek. It has 552B total parameters; 8B input activations, 16B output activations, and provides a context window of 1 million tokens. It is designed for fast inference and high-throughput tasks, delivering powerful reasoning and coding capabilities while maintaining excellent cost efficiency.

Its hybrid attention architecture enables efficient long-context processing. The model supports five inference levels—low, high, and max—with "medium" as the default and "max" representing the highest inference intensity. It is particularly well suited for coding assistants, conversational systems, and agentic workflows that require fast response times, scalability, and cost efficiency.

Base URL

https://api.icreat.ai/v1

Model Code

deepseek-v4.1-flash

Authentication

All API requests require authentication using an API key. You can obtain an API key from the console.

export ICREAT_API_KEY="your-api-key-here"

HTTP Request Headers

import os

API_KEY = os.environ.get("ICREAT_API_KEY")
headers = {
    "Content-Type": "application/json",
    "Authorization": "Bearer " + API_KEY,
}

Keep Your API Key Secure

Never expose your API key in client-side code or public repositories. Use environment variables or a backend proxy.

Code Examples

This model is accessed through an OpenAI-compatible Chat Completions API and supports both streaming and non-streaming modes. When stream is set to false (the default), the server returns the complete JSON response at once. When set to true, incremental content is delivered in chunks using Server-Sent Events (SSE), which is suitable for displaying the generation process in real time.

POST/chat/completions

Input Parameters

The following parameters are accepted in the request body.

Total: 7; Required: 2; Optional: 5

modelstringrequired

The model ID to use for the completion. You must specify the iCreat model_code (deepseek-v4.1-flash for this model), not the provider’s original model name.

Example: "deepseek-v4.1-flash"

messagesarray[object]required

The list of messages that make up the conversation.

max_tokensinteger

The maximum number of tokens to generate.

temperaturenumber

The sampling temperature, from 0 to 2.

streamboolean

When true, the response is streamed using Server-Sent Events.

reasoning_effortstring

Reasoning intensity. Supported values: high, xhigh (xhigh is the highest).

highxhigh
thinkingobject

Extended thinking mode configuration, if supported by the model.

Output Parameters

The API returns a response compatible with the OpenAI Chat Completions format.

Total: 6

idstring

The unique identifier for the completion.

objectstring

The object type, always chat.completion.

createdinteger

The Unix timestamp indicating when the completion was created.

modelstring

The model ID used for the completion.

choicesarray[object]

The list of completion choices.

usageobject

Token usage statistics.

LLM-Friendly Prompt

The following is an LLM-friendly Markdown prompt that can be copied into AI assistants such as Cursor or ChatGPT to help the AI understand the model’s API integration method, invocation workflow, and key parameters. Click the “Copy LLM Prompt” button or copy the full text from the code block below to copy the entire prompt.

# deepseek-v4.1-flash

> DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts (MoE) model developed by DeepSeek. It has 284 billion total parameters, activates 13 billion parameters per token, and provides a context window of 1 million tokens.

## Overview

Use iCreat’s OpenAI-compatible Chat Completions API for conversational completions, with support for streaming, non-streaming, and extended thinking modes.

## API Information

- **Base URL**: `https://api.icreat.ai/v1`
- **Endpoint (POST)**: `/chat/completions`
- **Model ID**: `deepseek-v4.1-flash`
- **Authentication**: `Authorization: Bearer ${ICREAT_API_KEY}`

## Invocation Workflow

Send a single POST request. With `stream: false`, the API returns a complete JSON response. With `stream: true`, the response is streamed using SSE.

### Input Notes

- `model` (required): The iCreat `model_code`, `deepseek-v4.1-flash`
- `messages` (required): The list of conversation messages
- Common optional fields: `max_tokens` (maximum output tokens), `temperature` (sampling temperature), `stream` (streaming toggle), and `thinking` (extended thinking mode)

### Output Notes

- Read the response from `choices[0].message.content`.

## Important Notes

- `model` must use the iCreat `model_code`.
- All other fields follow the OpenAI Chat Completions protocol.