• About Us
  • Privacy Policy
  • Disclaimer
  • Contact Us
AimactGrow
  • Home
  • Technology
  • AI
  • SEO
  • Coding
  • Gaming
  • Cybersecurity
  • Digital marketing
No Result
View All Result
  • Home
  • Technology
  • AI
  • SEO
  • Coding
  • Gaming
  • Cybersecurity
  • Digital marketing
No Result
View All Result
AimactGrow
No Result
View All Result

Baseten Provides DeepSeek-V4.1-Flash to Mannequin APIs With 1M-Token Context – Unite.AI

Admin by Admin
September 12, 2026
Home AI
Share on FacebookShare on Twitter



DeepSeek-V4.1-Flash is offered now on Baseten Mannequin APIs, Baseten introduced on September 11, 2026, bringing the 552B-parameter multimodal mixture-of-experts (MoE) mannequin, which pairs 8B energetic parameters for prefill with 16B for decode throughout a 1M-token context window, to the inference supplier’s platform.

DeepSeek launched the mannequin’s open weights on Hugging Face, and DeepSeek’s personal announcement is dated September 9, 2026. The mannequin accepts textual content and picture enter and generates textual content output, and the mannequin card states that the repository and weights are licensed underneath the MIT License. Baseten describes V4.1-Flash as DeepSeek’s third open-weight flash launch of 2026 and because the solely mannequin of its scale to make use of what DeepSeek calls a Causal Encoder-Decoder structure. Assist for Baseten’s Loops coaching product is coming quickly, in keeping with the corporate.

Reported Benchmark Outcomes

The mannequin card reviews instruct-model outcomes on the most reasoning effort setting of 100: V4.1-Flash scores 90.6 on Terminal-Bench 2.1, in contrast with 82.7 for V4-Flash and 87.9 for V4-Professional; 74.2 on DeepSWE v1.1, in contrast with 54.4 and 62.7; and 54.8 on AutomationBench, in contrast with 37.7 and 43.2. Baseten highlighted the identical coding and agentic figures, saying V4.1-Flash beats V4-Professional with roughly a 3rd of the entire parameters, whereas cautioning {that a} 54.8 on AutomationBench means the mannequin fails roughly half of advanced workflows and advising groups to maintain a human within the loop for agent pipelines.

Within the card’s comparability with frontier fashions at most effort, V4.1-Flash posts 90.9 on GPQA Diamond, a Codeforces score of 3471, and 63.9 on HLE with instruments. Baseten states that V4.1-Flash is DeepSeek’s first non-experimental mannequin with native picture enter, a functionality beforehand restricted to the experimental V4-Flash-Imaginative and prescient-Exp; its desk reviews 78.9 on Chartography and 49 on ZeroBench for the brand new mannequin, towards 64.3 and 35 for the experimental one.

Causal Encoder-Decoder Structure

In response to the mannequin card, V4.1-Flash organizes a 40-layer Transformer as a 20-layer causal encoder adopted by a 20-layer decoder, with the decoder’s international key-value (KV) cache projected from the ultimate encoder hidden states somewhat than derived from every decoder layer’s personal hidden states. The design prompts 8B parameters per token throughout prefill and 16B throughout decode; Baseten contrasts that with V4-Flash, which prompts 13B for each steps, framing the change as buying and selling a heavier decode for a a lot lighter prefill, a setup Baseten mentioned boosts value effectivity for coding brokers whose agentic loops generate way more prefill tokens than decode tokens.

The cardboard reviews that the mannequin’s Compressed Sparse Consideration 2 assigns every consideration layer certainly one of three static modes (Full, Reindex, or Reuse), with a Hierarchical Sparse Indexer within the decoder bounding deeper indexing value independently of context size. Mixed with FP4 essential KV caching, these designs scale back the worldwide KV cache to 890 bytes per token, roughly one quarter of V4-Flash, the cardboard states. A separate mechanism, SWA Bounded Replay, reconstructs lacking sliding-window-attention KV states by replaying solely the newest tokens, lowering the persistent KV footprint to roughly one eighth of V4-Flash. DeepSeek’s announcement places the financial savings at one quarter the HBM and one eighth the SSD storage of the earlier era.

Every MoE layer makes use of one shared knowledgeable and 384 routed specialists with six routed specialists energetic per token, and the mannequin provides Engram conditional reminiscence with 196B parameters alongside DSpark speculative decoding, in keeping with the cardboard. DeepSeek educated the mannequin from scratch on a 45T-token multimodal corpus, educated its sparse consideration at a 64K sequence size, and prolonged context to 1M tokens at 34T tokens. Submit-training follows an ordinary supervised fine-tuning, reinforcement studying, and on-policy distillation sequence, with substantive modifications concentrated in large-scale automated synthesis of agent duties and environments, and the mannequin exposes a repeatedly controllable reasoning effort setting from 1 to 100.

DeepSeek API Transition and Baseten Serving

DeepSeek states that V4-Flash and V4-Flash-Imaginative and prescient-Exp are retired on its platform, with the outdated API mannequin names briefly routing to V4.1-Flash for compatibility. New API pricing took impact at 04:00 UTC on September 10, 2026, with off-peak charges set at 50% of peak charges, and DeepSeek names official companions WorkBuddy (together with CodeBuddy) and OpenCode as totally supporting V4.1-Flash.

Baseten mentioned its Inference Stack serves the mannequin utilizing NVIDIA Dynamo with KV cache-aware routing, steering every request to the reproduction already holding its prefix somewhat than whichever reproduction is free. The mannequin is obtainable by way of Baseten’s Mannequin Library, with devoted deployments obtainable for groups needing reserved capability.

Beginning at 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests will path to V4.1-Flash at V4.1-Flash charges, an association DeepSeek mentioned will proceed till V4.1-Professional launches; the lab mentioned exams by a number of events put V4.1-Flash forward of V4-Professional on efficiency, value, velocity, and complete runtime.

Tags: 1MTokenaddsAPIsBasetenContextDeepSeekV4.1FlashmodelUnite.AI
Admin

Admin

Next Post
11 social media tendencies each marketer ought to watch in 2026 [new data]

11 social media tendencies each marketer ought to watch in 2026 [new data]

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Recommended.

Frequency Breathwork: Translating the Invisible Rhythm of Breath into Digital Type

Frequency Breathwork: Translating the Invisible Rhythm of Breath into Digital Type

December 29, 2025
Prime 21 Developer Newsletters to Subscribe To in 2025 — SitePoint

Prime 21 Developer Newsletters to Subscribe To in 2025 — SitePoint

April 24, 2025

Trending.

AI & data-driven Starbucks – Deep Brew

AI & data-driven Starbucks – Deep Brew

May 18, 2026
Self-Coding AI: Breakthrough or Hazard?

Self-Coding AI: Breakthrough or Hazard?

July 4, 2025
The Full Information to EcoGPT

The Full Information to EcoGPT

June 6, 2026
High LLM Observability and Analysis Platforms in 2026: Langfuse, LangSmith, Braintrust, Arize, and Extra In contrast

High LLM Observability and Analysis Platforms in 2026: Langfuse, LangSmith, Braintrust, Arize, and Extra In contrast

August 9, 2026
Hasbro Information Breach Uncovered Worker Private Data

Hasbro Information Breach Uncovered Worker Private Data

August 30, 2026

AimactGrow

Welcome to AimactGrow, your ultimate source for all things technology! Our mission is to provide insightful, up-to-date content on the latest advancements in technology, coding, gaming, digital marketing, SEO, cybersecurity, and artificial intelligence (AI).

Categories

  • AI
  • Coding
  • Cybersecurity
  • Digital marketing
  • Gaming
  • SEO
  • Technology

Recent News

11 social media tendencies each marketer ought to watch in 2026 [new data]

11 social media tendencies each marketer ought to watch in 2026 [new data]

September 12, 2026
Baseten Provides DeepSeek-V4.1-Flash to Mannequin APIs With 1M-Token Context – Unite.AI

Baseten Provides DeepSeek-V4.1-Flash to Mannequin APIs With 1M-Token Context – Unite.AI

September 12, 2026
  • About Us
  • Privacy Policy
  • Disclaimer
  • Contact Us

© 2025 https://blog.aimactgrow.com/ - All Rights Reserved

No Result
View All Result
  • Home
  • Technology
  • AI
  • SEO
  • Coding
  • Gaming
  • Cybersecurity
  • Digital marketing

© 2025 https://blog.aimactgrow.com/ - All Rights Reserved