model roundup

DeepSeek 4

3 items · started 2026-07-23 · closed 2026-07-26

  1. https://t.co/syBl6fkTDJ Miguel Salinas@VercantezHow we self-host DeepSeek V4 Flash on AWS spot instances4:18 PM · Jul 23, 20269.7KViews243560

  2. [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" a quiet day lets us highlight a new neolab win. Reignited distillation wars conversation aside, today was more of the same of previous news cycles, which…

  3. Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficie…

← all threads