Runwhere.AIRunwhere.AI
runwhere CLI

快速开始

用 runwhere 在 5 分钟内完成登录、同步、提交和查看日志。

这一页只保留最短路径。你先跑通一条任务链路,再回到其他页面看 YAML 细节、价格和完整命令。

1. 安装

ℹ️ 需要 Python 3.11 或更高版本。若系统 Python 过旧(如 macOS 自带 3.9),可用 uv 安装:uv tool install --python 3.12 runwhere。

pip install runwhere

开发环境可以在源码目录安装:

cd runwhere-runw
pip install -e ".[dev]"

确认 CLI 可用:

runwhere --help

2. 登录

runwhere login --endpoint https://api.runwhere.cn

命令会提示粘贴 API Token。非交互环境可以这样登录:

echo "$RUNWHERE_TOKEN" | runwhere login --endpoint https://api.runwhere.cn

验证身份:

runwhere whoami

3. 写一个最小 YAML

kind: Training
version: v1
job:
  name: hello

environment:
  image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
  # 装自己的依赖(可选,二选一;路径相对项目根,文件随 runwhere sync 上传)
  # requirements: ./requirements.txt
  # conda: myenv                # 项目根放 environment.yml
  command: python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

resources:
  gpuNum: 1
  # 机型指定:gpuSkuKey 填 8 位 SKU 编码(算力比价页「复制 SKU」)
  gpuSkuKey: PAPTQMKR

storage:
  workdirs:
    - path: "."

command 就是 GPU 世界的 hello world:不需要任何本地脚本,任务日志里打出 True NVIDIA T4 之类的输出,说明镜像、CUDA 和 GPU 整条链路都通了。跑通后再把 command 换成你自己的训练入口即可。

第一个任务建议在比价页挑低成本卡(t4 档位)的 SKU 试水,跑通链路后再换大卡 SKU 跑正式训练。

机型是怎么选的

用 gpuSkuKey 指定具体机型——这是可靠路径:

resources:
  gpuNum: 1
  gpuSkuKey: aliyun/cn-hangzhou//on_demand/ecs.gn7i-c4g1.xlarge

SKU KEY 的来源:控制台「算力比价」页每行有复制 SKU 按钮(示例里的值要换成你复制的真实 SKU)。

gpuType(如 t4)是便捷写法:提交时经比价服务在同型号里自动挑可用且便宜的。但依赖比价库有该型号库存——查不到时提交会被拒并提示改填 gpuSkuKey。稳妥起见,首选 gpuSkuKey。

需要 conda 吗

不需要。environment.image 指定的容器镜像就是运行环境——示例用的 pytorch 官方镜像自带 CUDA 和 PyTorch。

要装自己的依赖时,用 requirements(pip)或 conda 二选一,文件放在项目里随 runwhere sync 一起上传,平台在任务启动前自动构建并缓存(下次起任务直接复用):

pip 依赖 —— 文件放在 workdir 内(路径相对 workdir 根),如 requirements.txt:

transformers==4.44.2
datasets==2.20.0

YAML 里声明路径:

environment:
  image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
  requirements: requirements.txt
  command: python train.py

conda 环境 —— 项目根目录放 environment.yml(平台按 conda-lock.yml → conda-linux-64.lock → environment.yml → environment.yaml 顺序自动识别),conda 字段给环境起个名:

environment:
  image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
  conda: myenv
  command: python train.py

两个注意点:

  • requirements 和 conda 不能同时指定
  • conda 跨平台偶尔解不动(本机能装的包 Linux 上没有)——用 conda-lock -p linux-64 -f environment.yml 生成锁文件放进项目,平台会优先用这份精确锁

更多字段(args / env 等)见任务 YAML 的常用模板。

4. 同步

runwhere sync -f train.yaml

sync 会上传变更文件,并根据 YAML 准备环境和模型。

5. 提交

runwhere submit -f train.yaml

如果只写了 gpuType,CLI 会自动选择合适 Offer。生产任务建议先运行:

runwhere price recommend -f train.yaml --top 3

再把选中的 gpuSkuKey 写回 YAML。

6. 看日志

runwhere logs hello -f

任务结束或不用时:

runwhere stop job hello
runwhere delete job hello

下一步读什么

On this page