smoothquant

Here are 2 public repositories matching this topic...

intel / neural-compressor

Provide unified APIs for SOTA model compression techniques, such as low precision (INT8/INT4/FP4/NF4) quantization, sparsity, pruning, and knowledge distillation on mainstream AI frameworks such as TensorFlow, PyTorch, and ONNX Runtime.

sparsity pruning quantization knowledge-distillation auto-tuning low-precision quantization-aware-training post-training-quantization large-language-models smoothquant

Updated Nov 25, 2023
Python

intel / intel-extension-for-transformers

Star

⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡

chatbot stable-diffusion large-language-model chatpdf llm-inference smoothquant 4-bits speculative-decoding llm-cpu streamingllm attention-sink intel-optimized-llamacpp neural-chat

Updated Nov 25, 2023
C++

Improve this page

Add a description, image, and links to the smoothquant topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the smoothquant topic, visit your repo's landing page and select "manage topics."

Learn more