Qwen2模型量化时关于bitsandbytes安装的问题

时间：2024-09-19 14:53:45浏览次数：16

标签：bnb Qwen2 ids bitsandbytes input 量化 model True

Qwen2模型量化时关于bitsandbytes安装的问题

问题描述：

from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig,BitsAndBytesConfig
CUDA_DEVICE = "cuda:0"
model_name_or_path = '/qwen2-1.5b-instruct'
Tokenizer = AutoTokenizer.from_pretrained(model_name_or_path,use_fast=False)
bnb_config = BitsAndBytesConfig(
                    load_in_4bit=True,
                    bnb_4bit_use_double_quant=True,
                    bnb_4bit_quant_type="nf4",
                    bnb_4bit_compute_dtype=torch.bfloat16
                    )
Model = AutoModelForCausalLM.from_pretrained(model_name_or_path, 
												  device_map="auto",#CUDA_DEVICE, 
												  #load_in_8bit=True,
												  quantization_config=bnb_config,
												  #torch_dtype="auto"
												  )
messages = [
                {"role": "system", "content": "You are a helpful assistant."},
                {"role": "user", "content": prompt}
        ]

# "genarate chat" 
input_ids = Tokenizer.apply_chat_template(messages,tokenize=False,add_generation_prompt=True)
model_inputs = Tokenizer([input_ids], return_tensors="pt").to(CUDA_DEVICE)
generated_ids = Model.generate(model_inputs.input_ids,top_p=0.2,max_new_tokens=512)
generated_ids = [
	output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = self.Tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]

在模型量化时遇到如下问题：

RuntimeError:
        CUDA Setup failed despite GPU being available. Please run the following command to get mor                                                                                   e information:

        python -m bitsandbytes

        Inspect the output of the command and see if you can locate CUDA libraries. You might need                                                                                    to add them
        to your LD_LIBRARY_PATH. If you suspect a bug, please take the information from python -m                                                                                    bitsandbytes
        and open an issue at: https://github.com/TimDettmers/bitsandbytes/issues

解决方法1：

可以卸载重新安装bitsandbytes，
step1：pip uninstall bitsandbytes
step2：pip install bitsandbytes

python3 -m bitsandbytes

解决方法2：

Step1:确认系统查找动态库libcudart.so的路径,将查找到的路径添加进去便解决这个问题
find / -name “libcudart.so*” #我的libcudart.so在/usr/local/cuda-11.4中
export LD_LIBRARY_PATH=“/usr/local/cuda-11.4/bin:$LD_LIBRARY_PATH”
python3 -m bitsandbytes

参考资料：

《Ubuntu18.04+CUDA11安装bitsandbytes出现的问题》
https://blog.csdn.net/steptoward/article/details/135507131
《大模型训练时关于bitsandbytes安装的问题》
https://gitcode.csdn.net/662f78a69ab37021bfb31bf0.html
《bitsandbytes 报错》
https://blog.csdn.net/weixin_43967256/article/details/134006312

标签：bnb,Qwen2,ids,bitsandbytes,input,量化,model,True
From： https://blog.csdn.net/MITA1/article/details/142256842

只会Python编程，做量化交易策略用QMT怎么样？听说QMT是支持Python的！
QMT是专门为机构、活跃投资者、高净值客户等专业投资者研发的智能量化交易终端，拥有高速行情、极速交易、策略交易、多维度风控等专业功能，满足专业投资者的特殊交易需求。覆盖业务范围广:沪深A股、港股通、两融、期权、期货。适合用QMT的投资者：机构投资者:对系统交易工具和交......
可转债量化策略研究，QMT如何获取可转债合约信息？
获取可转债合约信息此函数被设计为专门用于单一转债的查询，能够提供详尽的转债信息。通过使用这个函数，您可以获取到深度的特定转债数据，包括其涨跌停价格、上市日期、退市日期和期权到期日等关键信息。这种全面的信息将成为您理解和分析转债历史趋势以及当前状态的有力工具。调......
设计能力量化
参考《体验设计案例课》，做记录。通用能力专业技能组织建设工作价值成本价值：从项目成本视角出发，提取数据来证明设计价值，比如说设计对流程的优化、降低了多少项目的人力成本、缩短了多长时间的上线周期；用户价值：提取一些用户视角的数据来证明设计价值，比如说获取用......
Qwen2-VL环境搭建&推理测试
引子2024年8月30号，阿里推出Qwen2-VL，开源了2B/7B模型，处理任意分辨率图像无需分割成块。之前写了一篇Qwen-VL的博客，感兴趣的童鞋请移步（Qwen-VL环境搭建&推理测试-CSDN博客），这么小的模型，显然我的机器是跑的起来的，OK，那就让我们开始吧。一、模型介绍Qwen2-VL的一项关键架构改进是......
高比例可再生能源电力系统的调峰成本量化与分摊模型（Matlab代码实现）
......
高比例可再生能源电力系统的调峰成本量化与分摊模型（Matlab代码实现）
......

Qwen2模型量化时关于bitsandbytes安装的问题