Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1631

Cascadia runs resident Inkling 975B/41B MoE inference on eleven Intel Core Ultra X7 AI PCs

The paper presents Cascadia’s custom resident MoE engine that executes Inkling (975B total, 41B active parameters) across eleven Intel Core Ultra X7 358H AI PCs with Arc B390 iGPU using OpenVINO. The system fits six decoder layers per machine, coordinates FP16 expert computation with FP32 output restoration, and reports up to 60.29 aggregate decode tokens/s at 88 streams (46.87 tokens/s over complete serving phases) and a median first-token latency of 6.05 s at fifteen streams; context lengths up to 64k tokens were evaluated, with first-token time growing as aN+bN^2 and decoding time scaling roughly linearly.

KEY POINTS

  1. The paper presents Cascadia’s custom resident MoE engine that executes Inkling (975B total, 41B active parameters) across eleven Intel Core Ultra X7 358H AI PCs with Arc B390 iGPU using OpenVINO.
  2. The system fits six decoder layers per machine, coordinates FP16 expert computation with FP32 output restoration, and reports up to 60.29 aggregate decode tokens/s at 88 streams (46.87 tokens/s over complete serving phases) and a median first-token latency of 6.05 s at fifteen streams; context lengths up to 64k tokens were evaluated, with first-token time growing as aN+bN^2 and decoding time scaling roughly linearly.
  3. This demonstrates a practical execution and evaluation approach for very large sparse MoE models on distributed client-grade PCs with shared CPU–GPU memory, widening possible deployment architectures beyond datacenters.

WHY IT MATTERS

This demonstrates a practical execution and evaluation approach for very large sparse MoE models on distributed client-grade PCs with shared CPU–GPU memory, widening possible deployment architectures beyond datacenters.

SOURCES & TIMELINE

1