Cold 8K prefill
Cold 8K prefill
8,192-token direct benchmark; median of three runs, Sep 7.
8,192 토큰 직접 벤치마크. 9월 7일, 3회 중앙값.
MODEL ENGINEERING / SSD-PLE
A large model becomes useful when the weights, memory layout, and runtime are designed together.
가중치, 메모리 배치, 런타임을 함께 설계해 대규모 모델을 실제로 활용합니다.
Qwen3.8 Flash Next combines a 128.8B-parameter compute backbone with a 51.2B-parameter predictive latent embedding (PLE) table. I packaged the table as SSD-backed sparse lookup memory, leaving accelerator memory for compute weights, KV state, and workspaces.
Qwen3.8 Flash Next는 128.8B 파라미터 연산 백본과 51.2B 파라미터 predictive latent embedding(PLE) 테이블을 결합합니다. PLE를 SSD 기반 sparse lookup으로 패키징해 연산 가중치, KV 상태, 워크스페이스를 위한 메모리를 확보했습니다.
Q5 main files: 77.56 GiB
BF16: 95.37 GiB · FP8: 47.68 GiB
These are weight-file sizes, not total running memory. The release requires the dedicated ds4 SSD-PLE loader; the GGUF format alone does not establish compatibility with other runtimes.
표시된 크기는 가중치 파일 크기이며 전체 실행 메모리가 아닙니다. 전용 ds4 SSD-PLE 로더가 필요하며 GGUF 형식 자체가 다른 런타임과의 호환성을 보장하지는 않습니다.
The work connects tensor-role quantization, SSD-PLE loading and caching, embedded MTP, image input, and session/KV handling. I directed profiling and bounded implementation experiments with coding agents, then reviewed output correctness and measured changes before retaining them.
텐서 역할별 양자화, SSD-PLE 로딩과 캐시, 내장 MTP, 이미지 입력, 세션·KV 처리를 연결했습니다. 코딩 에이전트와 프로파일링·구현 실험을 진행하고 출력 정확성과 측정 결과를 검토해 변경을 반영했습니다.
THE OPTIMIZATION JOURNEY
The first fixed 8,192-token direct prefill measured 145.02 tok/s. Q5_0 down-tail tiling, Gated DeltaNet state-column parallelization, and a bounded routed-MMQ worklist raised the same-command, same-input mean to 287.52 tok/s (+98.3%) on August 27.
첫 고정 8,192 토큰 직접 prefill은 145.02 tok/s였습니다. Q5_0 down-tail 타일링, Gated DeltaNet 상태 열 병렬화, bounded routed-MMQ worklist를 적용한 뒤 8월 27일 같은 명령·입력의 평균은 287.52 tok/s(+98.3%)가 됐습니다.
This is a development timeline. Later runs change the runtime, cache setup, or workload; the endpoints are not a controlled speedup ratio. Prefill and decode are reported separately.
개발 진행을 보여주는 타임라인입니다. 이후 측정은 런타임·캐시 설정·워크로드가 달라 시작과 끝을 통제된 배속 비교로 볼 수 없습니다. prefill과 decode는 구분해 보고합니다.
Further rounds fused expert-down work, removed repeated activation quantization and memory passes, overlapped PLE reads with compute, and reduced QSA synchronization. I used profiling to choose experiments and retained measured improvements after numerical and output checks.
이후 expert-down 연산 융합, 반복 activation 양자화와 메모리 패스 제거, PLE 읽기와 연산 중첩, QSA 동기화 감소를 진행했습니다. 프로파일링으로 실험을 정하고 수치·출력 검증 후 측정된 개선을 반영했습니다.
Initial benchmark and optimization record ↗GB10 · 128 GB unified memory · Q5 compute weights. The following are separate workloads, with their own protocols.
GB10 · 128 GB 통합 메모리 · Q5 연산 가중치. 아래 수치는 서로 다른 워크로드와 측정 조건의 결과입니다.
8,192-token direct benchmark; median of three runs, Sep 7.
8,192 토큰 직접 벤치마크. 9월 7일, 3회 중앙값.
8,036-token x-prompt; three fresh workers, two banks, MTP draft 2.
8,036 토큰 x-prompt. 새 worker 3개, 2 banks, MTP draft 2.
Incremental 2K–64K sweep; BF16 PLE, 2 GiB cache, MTP draft 2.
2K–64K 증분 sweep. BF16 PLE, 2 GiB 캐시, MTP draft 2.
BENCHMARK / ORIGINAL MODEL-CARD FIGURE
2,048-token incremental prefill and 128 greedy output tokens per frontier. Curves are per-frontier medians; bands show observed minimum and maximum over three fresh runs.
단계별 2,048 토큰 증분 prefill과 greedy 출력 128 토큰을 사용했습니다. 곡선은 단계별 중앙값, 음영은 새 프로세스 3회 실행의 관측 최솟값·최댓값입니다.
The September 8 paired comparison keeps the main Q5 GGUF and runtime fixed while changing only the PLE sidecar. The median of three run means over the 2K–64K sweep was 1,245.1 → 1,323.1 tok/s prefill (+6.3%) and 28.59 → 28.93 tok/s decode (+1.2%).
9월 8일 비교에서는 Q5 GGUF와 런타임을 고정하고 PLE 사이드카만 교체했습니다. 2K–64K sweep의 실행별 평균을 3회 측정한 중앙값은 prefill 1,245.1 → 1,323.1 tok/s(+6.3%), decode 28.59 → 28.93 tok/s(+1.2%)입니다.
Runs used interleaved fresh processes, a 2 GiB PLE cache, 16 workers, and MTP draft 2. Decode differences include generation and MTP policy effects. The smaller sidecar is not a claim of identical model quality.
새 프로세스를 교차 실행하고 PLE 캐시 2 GiB, worker 16개, MTP draft 2를 사용했습니다. decode 차이에는 생성 결과와 MTP 정책 영향이 포함되며, 사이드카 크기 감소가 모델 품질의 동일성을 의미하지는 않습니다.
FP8 sidecar protocol, raw CSV & limits ↗The base SSD-PLE repository recorded 17,140 Hugging Face downloads in the preceding month as of September 8, 2026. The model card publishes setup instructions, checksums, precision recipes, and dated measurement reports.
2026년 9월 8일 기준 기본 SSD-PLE 저장소의 Hugging Face 최근 한 달 다운로드는 17,140회입니다. 모델카드에 설정 방법, 체크섬, 정밀도 레시피, 날짜별 측정 리포트를 공개합니다.
Read the complete model card ↗