1、1大语言模型的异构计算和加速戴金权(Jason Dai)英特尔院士23自回归大语言模型(基于Transformer解码器架构)Transformer解码器架构自回归大语言模型:预测下一个token4Transformer解码器架构训练;推理(第一个token/Prefill)5Transformer解码器架构推理(下一个token/Decode)6大语言模型推理和训练瓶颈内存带宽计算显存大小分布式计算(互联)7大模型的异构计算和加速XPU异构计算 CPU,GPU,NPU硬件加速210 x Arc A770 GPU(16GB)客户端(Intel Core Ultra AI PC)边缘端(Inte
2、l AI座舱)服务器(Intel Xeon+Intel Arc GPUs)8大模型的异构计算和加速低比特计算 模型量化/压缩(WxAy)数据类型(INTx,FPx)低比特算子 显存(如kv cache)使用量 训练、微调(如QLoRA)9低比特大模型的精度困惑度(Wikitext数据集)10大模型的异构计算和加速推理算法优化 Self-speculative decoding KV Cache compression Sliding window attention Sparse attention Flash attention/decoding Continuous batching Pr
3、efill/decoding disaggregation 11IPEX-LLM:开源大模型XPU加速框架XPU ComputePython(PyTorch)Ecosystemllama.cpp EcosystemHuggingFace,Langchain,LlamaIndex,DeepSpeed,TRL,Axolotl,IPEX-LLM LibraryUsers/DevelopersIntel XPUllama.cpp,Ollama,LangChain.js,Open WebUI,LLM Accelerationhttps:/ XPU 大模型加速体验13Intel UHD/Iris iGPU
4、llama.cpp+IPEX-LLM(Phi-3-mini,Q4_0)14Intel Core Ultra AI PC Ollama+IPEX-LLM(Mistral-7B,Q4_K_M)15Intel Arc A770 GPUTextGeneration-WebUI+IPEX-LLM(Llama3-8B,FP8)164 x Arc A770 GPUFastChat+IPEX-LLM(QWen1.5-72B FP6)17LoRA/QLoRA on Xeon+Multi-Arc支持 PEFT,TRL,Axolotl,Zero2/Zero318英特尔 XPU 大模型应用创新19Office助手Ex
5、tendOffice展示20工业机器人代码生成科东软件展示21AI座舱-汽车助理智谱AI展示22AI座舱-驾驶伴侣百川智能展示23个人或企业本地RAG系统在英特尔 XPU 上运行 RAGFlow(https:/ XPU 上运行 GraphRAG(https:/ to Actions关注和试用 IPEX-LLM,并给我们反馈https:/ IPEX-LLM 在Intel XPU 平台开发大模型及其应用 客户端-边缘-服务器(Intel Core Ultra AI PC、AI座舱、Xeon+Intel Arc GPUs)高效的大模型 XPU 加速的创新 大模型应用场景的创新25谢谢!2627Not
6、ices&DisclaimersPerformance varies by use,configuration and other factors.Learn more on the Performance Index site.Performance results are based on testing as of dates shown in configurations and may not reflect all publicly available updat