快速开始
用 runwhere 在 5 分钟内完成登录、同步、提交和查看日志。
这一页只保留最短路径。你先跑通一条任务链路,再回到其他页面看 YAML 细节、价格和完整命令。
1. 安装
ℹ️ 需要 Python 3.11 或更高版本。若系统 Python 过旧(如 macOS 自带 3.9),可用 uv 安装:
uv tool install --python 3.12 runwhere。
pip install runwhere开发环境可以在源码目录安装:
cd runwhere-runw
pip install -e ".[dev]"确认 CLI 可用:
runwhere --help2. 登录
runwhere login --endpoint https://api.runwhere.cn命令会提示粘贴 API Token。非交互环境可以这样登录:
echo "$RUNWHERE_TOKEN" | runwhere login --endpoint https://api.runwhere.cn验证身份:
runwhere whoami3. 写一个最小 YAML
kind: Training
version: v1
job:
name: hello
environment:
image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
# 装自己的依赖(可选,二选一;路径相对项目根,文件随 runwhere sync 上传)
# requirements: ./requirements.txt
# conda: myenv # 项目根放 environment.yml
command: python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"
resources:
gpuNum: 1
# 机型指定:gpuSkuKey 填 8 位 SKU 编码(算力比价页「复制 SKU」)
gpuSkuKey: PAPTQMKR
storage:
workdirs:
- path: "."command 就是 GPU 世界的 hello world:不需要任何本地脚本,任务日志里打出 True NVIDIA T4 之类的输出,说明镜像、CUDA 和 GPU 整条链路都通了。跑通后再把 command 换成你自己的训练入口即可。
第一个任务建议在比价页挑低成本卡(t4 档位)的 SKU 试水,跑通链路后再换大卡 SKU 跑正式训练。
机型是怎么选的
用 gpuSkuKey 指定具体机型——这是可靠路径:
resources:
gpuNum: 1
gpuSkuKey: aliyun/cn-hangzhou//on_demand/ecs.gn7i-c4g1.xlargeSKU KEY 的来源:控制台「算力比价」页每行有复制 SKU 按钮(示例里的值要换成你复制的真实 SKU)。
gpuType(如 t4)是便捷写法:提交时经比价服务在同型号里自动挑可用且便宜的。但依赖比价库有该型号库存——查不到时提交会被拒并提示改填 gpuSkuKey。稳妥起见,首选 gpuSkuKey。
需要 conda 吗
不需要。environment.image 指定的容器镜像就是运行环境——示例用的 pytorch 官方镜像自带 CUDA 和 PyTorch。
要装自己的依赖时,用 requirements(pip)或 conda 二选一,文件放在项目里随 runwhere sync 一起上传,平台在任务启动前自动构建并缓存(下次起任务直接复用):
pip 依赖 —— 文件放在 workdir 内(路径相对 workdir 根),如 requirements.txt:
transformers==4.44.2
datasets==2.20.0YAML 里声明路径:
environment:
image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
requirements: requirements.txt
command: python train.pyconda 环境 —— 项目根目录放 environment.yml(平台按 conda-lock.yml → conda-linux-64.lock → environment.yml → environment.yaml 顺序自动识别),conda 字段给环境起个名:
environment:
image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
conda: myenv
command: python train.py两个注意点:
requirements和conda不能同时指定- conda 跨平台偶尔解不动(本机能装的包 Linux 上没有)——用
conda-lock -p linux-64 -f environment.yml生成锁文件放进项目,平台会优先用这份精确锁
更多字段(args / env 等)见任务 YAML 的常用模板。
4. 同步
runwhere sync -f train.yamlsync 会上传变更文件,并根据 YAML 准备环境和模型。
5. 提交
runwhere submit -f train.yaml如果只写了 gpuType,CLI 会自动选择合适 Offer。生产任务建议先运行:
runwhere price recommend -f train.yaml --top 3再把选中的 gpuSkuKey 写回 YAML。
6. 看日志
runwhere logs hello -f任务结束或不用时:
runwhere stop job hello
runwhere delete job hello