Sylvia Xiao | AI Engineer
SylviaXiao
文章
14
分类
2
标签
16
Lazy loaded image
Read Blog
🏋🏼Memory and Computation in LLM Inference: Transformers vs. RWKV
发布于: 2025-2-23
最后更新: 2026-5-28
次查看
llm
RWKV
Tansformer
目录
0%
Memory and Computation in LLM Inference: Transformers vs. RWKV1. Memory Usage and GPU Utilization1.1 Memory Usage (VRAM)1.2 GPU Utilization2. Data Flow in Transformer-Based LLMs2.1 Two-Stage LLM Inference with KV-Cache2.1.1 Prompt Stage2.1.2 Decoding Stage2.2 Self-Attention and KV-Cache in Standard Transformers2.2.1 Computational Complexity Analysis2.2.2 Space Complexity3. The RWKV Architecture3.1 Data Flow in RWKV3.2 Time Mixing and Channel Mixing Details3.3 RWKV Inference Process3.3.1 Processing the Prompt3.3.2 Generating New Tokens3.4 Memory Usage in RWKV3.5 A Pitfall in the MLC Framework’s RWKV5 Implementation4. Comparison of RAM Usage: Transformers vs. RWKV
SylviaXiao
SylviaXiao
AI Engineer
文章
14
分类
2
标签
16
最新发布
Qwen-Image Text-to-Image Pipeline Note
Qwen-Image Text-to-Image Pipeline Note
2026-6-22
Whisper’s Encoder-Decoder Pipeline, Special Tokens, and a Comparison with Qwen3-ASR
Whisper’s Encoder-Decoder Pipeline, Special Tokens, and a Comparison with Qwen3-ASR
2026-6-17
Financial / Universe Robot Dialogue AI Interface
Financial / Universe Robot Dialogue AI Interface
2026-6-2
Real-time Logging in Model Training Using Python
Real-time Logging in Model Training Using Python
2026-5-28
Building MLC-LLM: A Step-by-Step Guide
Building MLC-LLM: A Step-by-Step Guide
2026-5-28
Memory and Computation in LLM Inference: Transformers vs. RWKV
Memory and Computation in LLM Inference: Transformers vs. RWKV
2026-5-28
目录
0%
Memory and Computation in LLM Inference: Transformers vs. RWKV1. Memory Usage and GPU Utilization1.1 Memory Usage (VRAM)1.2 GPU Utilization2. Data Flow in Transformer-Based LLMs2.1 Two-Stage LLM Inference with KV-Cache2.1.1 Prompt Stage2.1.2 Decoding Stage2.2 Self-Attention and KV-Cache in Standard Transformers2.2.1 Computational Complexity Analysis2.2.2 Space Complexity3. The RWKV Architecture3.1 Data Flow in RWKV3.2 Time Mixing and Channel Mixing Details3.3 RWKV Inference Process3.3.1 Processing the Prompt3.3.2 Generating New Tokens3.4 Memory Usage in RWKV3.5 A Pitfall in the MLC Framework’s RWKV5 Implementation4. Comparison of RAM Usage: Transformers vs. RWKV
2021-2026SylviaXiao.

SylviaXiao | AI Engineer | AI Engineer

Powered byNotionNext 4.10.10.