在个人电脑一键运行谷歌最新 Gemma-2-9B 大模型

文摘 2024-07-02 18:45 英国

谷歌最近发布了9B和27B大小的 Gemma 2模型^[1]，这是其 Gemma 模型系列的最新型号。根据其技术报告，未来几天将开源一个 Gemma-2-2b 模型。技术报告还显示，Gemma-2-9B模型在多个基准测试中的表现超过了 Mistral-7B、Llama-3-8B和 Gemma 1.5模型。

如果想一键在你的计算机上运行 Gemma-9b-Chat，可以在终端中运行以下命令
bash <(curl -sSfL 'https://raw.githubusercontent.com/LlamaEdge/LlamaEdge/main/run-llm.sh') —model gemma-2-9b-it

本文将以 Gemma-2-9B 为例，手把手教你轻松

在自己的设备上运行 Gemma-2-9B on your own device
为 Gemma-2-9B 创建一个 API server

我们将使用 LlamaEdge^[2] （Rust + Wasm技术栈）开发并部署该模型的应用。无需安装复杂的 Python 包或 C++ 工具链！了解我们为什么选择 Rust + Wasm^[3]。

在你自己的设备上运行 Gemma-2-9B

第一步：用下面的命令行安装 WasmEdge^[4].

curl -sSf https://raw.githubusercontent.com/WasmEdge/WasmEdge/master/utils/install_v2.sh | bash -s -- -v 0.13.5

第二步：下载 Gemma-2-9B-it model GGUF 模型^[5]文件。模型有6.40 GB，所以下载可能需要一定时间

curl -LO https://huggingface.co/second-state/gemma-2-9b-it-GGUF/resolve/main/gemma-2-9b-it-Q5_K_M.gguf

第三步：下载一个跨平台的可移植 Wasm 文件，是个聊天应用。该应用让你能在命令行中与模型聊天。该应用的 Rust 源代码戳这里^[6]。

curl -LO https://github.com/LlamaEdge/LlamaEdge/releases/latest/download/llama-chat.wasm

就是这样。可以通过输入以下命令在终端与模型聊天。

wasmedge --dir .:. --nn-preload default:GGML:AUTO:gemma-2-9b-it-Q5_K_M.gguf llama-chat.wasm -p gemma-instruct -c 4096

这个可移植的Wasm应用会自动利用设备上的硬件加速器（例如GPU）。

[You]:
Tell me what is gemma?

[Bot]:
Gemma is an open-weights AI assistant developed by Google DeepMind.

Here are some key things to know about Gemma:

* **Open Weights:** This means the underlying code and model weights are freely accessible to the public. Anyone can inspect, modify, or build upon Gemma.
* **Text-Only:** Gemma is designed primarily for text-based interactions. It can understand and generate human-like text but doesn't have capabilities like image recognition or voice synthesis.
* **Focus on Research & Development:** Gemma is intended as a tool for researchers and developers to explore and experiment with different AI models and techniques.

**Why is Gemma significant?**

* **Transparency and Collaboration:** Open-weights nature promotes transparency in AI development and encourages collaboration among researchers worldwide.
* **Customization:** Users can adapt Gemma to their specific needs by fine-tuning it on their own datasets.
* **Accessibility:**  Makes advanced AI technology accessible to a wider range of individuals and organizations, beyond those with large resources.

**Keep in mind:**

* Gemma is still under development and may have limitations compared to more mature, closed-source models.
* Its capabilities are primarily focused on text processing.

为Gemma-2-9b-it^[7] 创建一个兼容OpenAI的 API server

一个兼容 OpenAI 的API 使得 Llama-3-8B-Chinese 能够与不同的开发框架和工具无缝集成，比如 flows.network^[8], LangChain and LlamaIndex等等，提供更广泛的应用可能。大家也可以参考其代码自己写自己的API服务器或者其它大模型应用。想要启动 API 服务，请按以下步骤操作：下载这个 API 服务器应用。它是一个跨平台的可移植 Wasm 应用，可以在各种 CPU 和 GPU 设备上运行。

curl -LO https://github.com/LlamaEdge/LlamaEdge/releases/latest/download/llama-api-server.wasm

然后，下载聊天机器人 Web UI，从而通过聊天机器人 UI 与模型进行交互。

curl -LO https://github.com/LlamaEdge/chatbot-ui/releases/latest/download/chatbot-ui.tar.gz
tar xzf chatbot-ui.tar.gz
rm chatbot-ui.tar.gz

接下来，使用以下命令行启动模型的 API 服务器。然后，打开浏览器访问 http://localhost:8080^[9] 开始聊天！

wasmedge --dir .:. --nn-preload default:GGML:AUTO:gemma-2b-it-Q5_K_M.gguf llama-api-server.wasm -p gemma-instruct -c 4096

另外打开一个终端窗口，可以使用 curl 与 API 服务器进行交互。

curl -X POST http://localhost:8080/v1/chat/completions \
  -H 'accept:application/json' \
  -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"system", "content": "You are a sentient, superintelligent artificial general intelligence, here to teach and assist me."}, {"role":"user", "content": "Write a short story about Goku discovering kirby has teamed up with Majin Buu to destroy the world."}], "model":"Gemma-2b-it"}'

就是这样啦。WasmEdge 是运行 LLM 应用最简单、最快、最安全的方式^[10]。快来试试看吧！

参考资料

[1]

9B和27B大小的Gemma 2模型: https://ai.google.dev/gemma/docs

[2]

Image: image.png]我们将使用 [LlamaEdge: https://github.com/second-state/LlamaEdge/

[3]

了解我们为什么选择 Rust + Wasm: https://www.secondstate.io/articles/fast-llm-inference/

[4]

WasmEdge: https://github.com/WasmEdge/WasmEdge

[5]

Gemma-2-9B-it model GGUF 模型: https://huggingface.co/second-state/gemma-2-9b-it-GGUF

[6]

这里: https://github.com/second-state/llama-utils/tree/main/chat

[7]

Gemma-2-9b-it: https://www.secondstate.io/articles/gemma-2-9b/#create-an-openai-compatible-api-service-for-gemma-2-9b-it

[8]

flows.network: https://flows.network/

[9]

http://localhost:8080: http://localhost:8080/

[10]

运行 LLM 应用最简单、最快、最安全的方式: https://www.secondstate.io/articles/fast-llm-inference/

关于 WasmEdge

WasmEdge 是轻量级、安全、高性能、可扩展、兼容OCI的软件容器与运行环境。目前是 CNCF 沙箱项目。WasmEdge 被应用在 SaaS、云原生，service mesh、边缘计算、边缘云、微服务、流数据处理、LLM 推理等领域。

GitHub：https://github.com/WasmEdge/WasmEdge

官网：https://wasmedge.org/

‍‍Discord 群：https://discord.gg/U4B5sFTkFc

文档：https://wasmedge.org/docs

http://mp.weixin.qq.com/s?__biz=MzI2MjkxNjA2Mg==&mid=2247487196&idx=1&sn=261a391c3c6fcead2f4cac29d76d639a

Second State

Rust 函数即服务

在昇腾 910B 上部署轻量级和跨平台大模型 Agent

课程升级、资源加码！万人共学的书生大模型实战营第4期正式起航！

OSC源创会·北京站：高性能计算与大模型推理

RTE 大会报名丨AI 时代新基建：云边端架构和 AI Infra ，RTE2024 技术专场第二弹！

2024年第五届CID参会就在明天！

Rust 群星闪耀！20+ 海内外顶尖 Rust 天团 GOSIM CHINA 2024 相聚北京

开创跨平台的未来！GOSIM CHINA 2024《App 开发》专题论坛重磅揭晓！

打造更安全、去中心化和协作的互联网！GOSIM CHINA 2024《下一代互联网》重磅嘉宾揭晓

Triton & vLLM 联袂呈现 AI 技术盛宴：高效推理框架的应用实践与未来创新

倒计时 2 天，GOSIM CHINA 2024 全日程重磅发布（附参会指南）！

聚焦开源大模型前沿应用，GOSIM CHINA 2024《AI 模型与基础模型》专题论坛重磅揭晓！

ChatGPT开源替代：阿里最新最强大模型千问2.5

在 MacBook 上运行 FLUX.1，可无缝跨平台 | 为假期添加点趣味

Wasm技术浪潮来袭：加入我们的在线课程，掌握WebAssembly的未来

贡献开源拿奖励，再送10份免费课程/认证考试

自建AI编程助手 | 本地 Yi-Coder模型 + Cursor 5分钟写一个网页

议题征集倒计时啦！不能错过的第五届CID大会！

当 Rust 遇到 AI 会擦出什么样的火花|与你相约 RustChinaConf 2024

Mac上运行微软最新Phi-3.5-mini大模型+开发Agent

【福利】来偶遇Linus！KubeCon + CloudNativeCon +开源峰会+ AI_dev China下周三火热开幕

来 RustChinaConf 听听 LlamaEdge 的 Rust 实践

极客与技术，产业与生态，年度开源峰会 2024 GOTC x GOGC 即将开幕

2024 秋季WasmEdge LFX实习机会：大模型、交易机器人等你来

LlamaEdge 支持 tool call！调用外部工具

KubeCon 2024 AI_Dev日程已发布!

本地搭建 AI 服务？一文带你轻松部署 internlm2_5-7b-chat 大模型应用

在 Llama 3.1 构建多种AI应用

简单命令行搭建吴恩达的 LLM Translation Agent，测测开源模型哪家强

《歌手》排名里的 13.8%和13.11%哪个大？ Mathstral：AI数学能力大考验！

在个人电脑一键运行谷歌最新 Gemma-2-9B 大模型

OpenAI 不可用？使用开源模型一键替换 OpenAI API

扫码申请最终用户门票｜2024 年 KubeCon + CloudNativeCon + 开源峰会 + AI_dev 中国大会

阿里巴巴全球数学竞赛是什么难度？让阿里的Qwen2-72B 试一试

2024 年 KubeCon + CloudNativeCon + 开源峰会 + AI_dev 中国大会的精彩阵容出炉！

做大模型时代的开源贡献者，WasmEdge 开源之夏项目等你来

一键运行零一万物新鲜出炉Yi-1.5-9B-Chat大模型

Llama-3-8B 中文版来了，在自己设备上运行试试看吧

Wasm 性能究竟如何 | Arm 上的容器运行时和 WasmEdge 基准测试

本周末来上海 GOTC 现场和 WasmEdge 见面吧

Open Source Summit NA 上的 WebAssembly演讲

KubeCon EU |云计算的未来是什么？

开源之夏2023明天开启报名！欢迎报名 WasmEdge 社区项目

WebAssembly @ KubeCon + CloudNativeCon EU 2023

那些让 ChatGPT review 代码的程序员，后来都怎么样了？

用 Rust 开发 WasmEdge 应用 | 微软 Reactor 活动回顾

社区合作|第二届开源云原生开发者日开启预约！

五分钟创建一个 Serverless ChatGPT GitHub App

活动预告|【欧拉多咖·操作系统研讨会】第九期：面向未来云计算的虚拟化技术

分类

时事

民生

政务

教育

文化

科技

财富

体娱

健康

情感

旅行

百科

职场

楼市

企业

乐活

学术

汽车

时尚

创业

美食

幽默

美体

文摘

原创标签

时事社会财经军事教育体育科技汽车科学房产搞笑综艺明星音乐动漫游戏时尚健康旅游美食生活摄影宠物职场育儿情感小说曲艺文化历史三农文学娱乐电影视频图片新闻宗教电视剧纪录片广告创意壁纸头像心灵鸡汤星座命理教育培训艺术文化金融财经健康医疗美妆时尚餐饮美食母婴育儿社会新闻工业农业时事政治星座占卜幽默笑话独立短篇连载作品文化历史科技互联网

发布位置

广东北京山东江苏河南浙江山西福建河北上海四川陕西湖南安徽湖北内蒙古江西云南广西甘肃辽宁黑龙江贵州新疆重庆吉林天津海南青海宁夏西藏香港澳门台湾美国加拿大澳大利亚日本新加坡英国西班牙新西兰韩国泰国法国德国意大利缅甸菲律宾马来西亚越南荷兰柬埔寨俄罗斯巴西智利卢森堡芬兰瑞典比利时瑞士土耳其斐济挪威朝鲜尼日利亚阿根廷匈牙利爱尔兰印度老挝葡萄牙乌克兰印度尼西亚哈萨克斯坦塔吉克斯坦希腊南非蒙古奥地利肯尼亚加纳丹麦津巴布韦埃及坦桑尼亚捷克阿联酋安哥拉

在个人电脑一键运行谷歌最新 Gemma-2-9B 大模型

为Gemma-2-9b-it[7] 创建一个兼容OpenAI的 API server

为Gemma-2-9b-it^[7] 创建一个兼容OpenAI的 API server