Implement balance monitoring and notification for Selectel API integration
- Added support for balance tracking using `SELECTEL_API_TOKEN` in the configuration. - Introduced new balance notification logic in the local watchdog, alerting users on balance changes based on defined thresholds. - Updated documentation to include instructions for setting up balance notifications and the required environment variables. - Enhanced the `ready` and `session` modules to initialize balance state and handle notifications during GPU operations. - Refactored the CLI and related components to support the new balance monitoring features, ensuring a seamless user experience.
This commit is contained in:
+13
-2
@@ -231,8 +231,9 @@ Application credential для idle-killer CLI создаёт на `up` (узко
|
|||||||
|
|
||||||
1. Nova `ACTIVE`
|
1. Nova `ACTIVE`
|
||||||
2. TCP 22 / SSH
|
2. TCP 22 / SSH
|
||||||
3. HTTP `http://127.0.0.1:7801` **на VM** (через SSH)
|
3. На VM: HTTP сервисов стека (SwarmUI `:7801` / Ollama `:11434` / llama.cpp `:8080`) — `verify_stack_on_vm`
|
||||||
4. Backend Idle → toast (если `NOTIFY_READY`) + access-card
|
4. Backend Idle (если SwarmUI) → toast (если `NOTIFY_READY`)
|
||||||
|
5. Туннель + проверка **localhost** тех же сервисов → access-card
|
||||||
|
|
||||||
Локальный порт UI: **17801** (на VM по-прежнему 7801 на loopback).
|
Локальный порт UI: **17801** (на VM по-прежнему 7801 на loopback).
|
||||||
|
|
||||||
@@ -241,3 +242,13 @@ Application credential для idle-killer CLI создаёт на `up` (узко
|
|||||||
## Windows / notify
|
## Windows / notify
|
||||||
|
|
||||||
`NOTIFY_READY`: toast + системный звук, когда backend Idle. Если toast недоступен — только звук и лог, без падения CLI.
|
`NOTIFY_READY`: toast + системный звук, когда backend Idle. Если toast недоступен — только звук и лог, без падения CLI.
|
||||||
|
|
||||||
|
### Баланс Selectel (шаг 200 ₽)
|
||||||
|
|
||||||
|
Нужны:
|
||||||
|
|
||||||
|
1. `SELECTEL_API_TOKEN` в `.env` — **статический ключ панели** (`X-Token`), не OpenStack password ([баланс API](https://docs.selectel.ru/api/balance/)).
|
||||||
|
2. `gpu-rent watchdog install` — тот же локальный тикер, что и аварийный stop.
|
||||||
|
|
||||||
|
На `up` запоминается baseline баланса. Каждый tick watchdog: если с `up` ушло ещё ~200 ₽ (`BALANCE_NOTIFY_STEP_RUB`) — toast «Списано ~N ₽…». Пополнение сбрасывает baseline. Опционально `BALANCE_NOTIFY_LOW_RUB=300` — разовое «баланс низкий».
|
||||||
|
|
||||||
|
|||||||
@@ -26,6 +26,7 @@
|
|||||||
| llm-only | `ENABLE_SWARMUI=false` / `WORKLOAD=llm` / `--no-swarm`/`--llm-only`: GPU + LLM без SwarmUI (bootstrap только data disk) |
|
| llm-only | `ENABLE_SWARMUI=false` / `WORKLOAD=llm` / `--no-swarm`/`--llm-only`: GPU + LLM без SwarmUI (bootstrap только data disk) |
|
||||||
| Capture | `gpu-rent capture` — инвентарь VM → merge **ссылок** в локальные yaml (веса не качать) |
|
| Capture | `gpu-rent capture` — инвентарь VM → merge **ссылок** в локальные yaml (веса не качать) |
|
||||||
| Perf auto-tune | На `up`: probe GPU → tier. Swarm/Comfy: sageattention ExtraArgs на Ampere+ ≥16 GiB. Ollama: flash/KV/keep-alive + `GPU_OVERHEAD` чтобы оставить VRAM под Krea |
|
| Perf auto-tune | На `up`: probe GPU → tier. Swarm/Comfy: sageattention ExtraArgs на Ampere+ ≥16 GiB. Ollama: flash/KV/keep-alive + `GPU_OVERHEAD` чтобы оставить VRAM под Krea |
|
||||||
|
| Баланс ₽ | `SELECTEL_API_TOKEN` (X-Token панели) + `gpu-rent watchdog install`. Baseline на `up`; тик watchdog уведомляет каждые `BALANCE_NOTIFY_STEP_RUB` (дефолт **200**) списанных с baseline. Опционально `BALANCE_NOTIFY_LOW_RUB` |
|
||||||
| Data-диск | Старт **100 GB**, рост через resize вверх (вниз Selectel не умеет) |
|
| Data-диск | Старт **100 GB**, рост через resize вверх (вниз Selectel не умеет) |
|
||||||
| SSH | CLI генерирует `<repo>/.gpu-rent/id_ed25519` без passphrase и сам регистрирует keypair |
|
| SSH | CLI генерирует `<repo>/.gpu-rent/id_ed25519` без passphrase и сам регистрирует keypair |
|
||||||
| Локальные файлы | Всё в корне репозитория: `.env`, `models.yaml`, `extensions.yaml`; runtime (`state.json`, lock, SSH) в `<repo>/.gpu-rent/`. Не `%USERPROFILE%\.gpu-rent` |
|
| Локальные файлы | Всё в корне репозитория: `.env`, `models.yaml`, `extensions.yaml`; runtime (`state.json`, lock, SSH) в `<repo>/.gpu-rent/`. Не `%USERPROFILE%\.gpu-rent` |
|
||||||
|
|||||||
@@ -19,6 +19,8 @@ GPU есть не во всех сегментах. Перед create — [ма
|
|||||||
|
|
||||||
Статический API-ключ панели (`X-Token`) **не управляет** объектами OpenStack (серверы, диски, сети).
|
Статический API-ключ панели (`X-Token`) **не управляет** объектами OpenStack (серверы, диски, сети).
|
||||||
|
|
||||||
|
Исключение: **баланс аккаунта** (`GET https://api.selectel.ru/v3/balances`) — как раз через `X-Token`. Для уведомлений о списании положи `SELECTEL_API_TOKEN` в `.env` и поставь `gpu-rent watchdog install` (см. [cli.md](cli.md)).
|
||||||
|
|
||||||
Нужен **сервисный пользователь** с правом на проект.
|
Нужен **сервисный пользователь** с правом на проект.
|
||||||
|
|
||||||
1. CLI хранит: account id, username, password, project id/name, pool, pool segment.
|
1. CLI хранит: account id, username, password, project id/name, pool, pool segment.
|
||||||
|
|||||||
+1
-1
@@ -233,7 +233,7 @@ copy models.example.yaml models.yaml
|
|||||||
1. Короткий doctor (полный — `up -v`).
|
1. Короткий doctor (полный — `up -v`).
|
||||||
2. Create/unshelve GPU + диски (после confirm).
|
2. Create/unshelve GPU + диски (после confirm).
|
||||||
3. Bootstrap SwarmUI, extensions, autocomplete, Civitai-seed, push локальных папок; optional LLM.
|
3. Bootstrap SwarmUI, extensions, autocomplete, Civitai-seed, push локальных папок; optional LLM.
|
||||||
4. Туннель на `localhost:17801`, access-card с URL / MCP.
|
4. Туннель на `localhost:17801` (или LLM-порт), **проверка** что сервисы отвечают, access-card.
|
||||||
5. Процесс ждёт: **Ctrl+C** закрывает только туннель, GPU остаётся.
|
5. Процесс ждёт: **Ctrl+C** закрывает только туннель, GPU остаётся.
|
||||||
|
|
||||||
Полезные флаги:
|
Полезные флаги:
|
||||||
|
|||||||
@@ -63,3 +63,8 @@ IDLE_GRACE_MINUTES=45
|
|||||||
# LOCAL_WATCHDOG_GRACE_MINUTES=10
|
# LOCAL_WATCHDOG_GRACE_MINUTES=10
|
||||||
PULL_OUTPUT=false
|
PULL_OUTPUT=false
|
||||||
NOTIFY_READY=true
|
NOTIFY_READY=true
|
||||||
|
# Статический API-ключ панели (X-Token) — только для баланса, не для OpenStack:
|
||||||
|
# SELECTEL_API_TOKEN=
|
||||||
|
# Уведомления о списании (local-watchdog tick), шаг с момента up:
|
||||||
|
# BALANCE_NOTIFY_STEP_RUB=200
|
||||||
|
# BALANCE_NOTIFY_LOW_RUB=0
|
||||||
|
|||||||
@@ -0,0 +1,174 @@
|
|||||||
|
"""Selectel account balance watch: notify every N ₽ spent since up."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from collections.abc import Callable
|
||||||
|
from dataclasses import asdict, dataclass
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
from gpu_rent.errors import CloudError
|
||||||
|
from gpu_rent.paths import runtime_dir
|
||||||
|
|
||||||
|
Log = Callable[[str], None]
|
||||||
|
|
||||||
|
BALANCES_URL = "https://api.selectel.ru/v3/balances"
|
||||||
|
DEFAULT_STEP_RUB = 200.0
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class BalanceWatchState:
|
||||||
|
baseline_rub: float
|
||||||
|
last_notified_step: int = 0
|
||||||
|
low_notified: bool = False
|
||||||
|
last_balance_rub: float | None = None
|
||||||
|
updated_at: str | None = None
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return asdict(self)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def from_dict(cls, data: dict[str, Any]) -> BalanceWatchState:
|
||||||
|
return cls(
|
||||||
|
baseline_rub=float(data.get("baseline_rub") or 0),
|
||||||
|
last_notified_step=int(data.get("last_notified_step") or 0),
|
||||||
|
low_notified=bool(data.get("low_notified")),
|
||||||
|
last_balance_rub=(
|
||||||
|
float(data["last_balance_rub"])
|
||||||
|
if data.get("last_balance_rub") is not None
|
||||||
|
else None
|
||||||
|
),
|
||||||
|
updated_at=data.get("updated_at"),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def balance_watch_path() -> Path:
|
||||||
|
return runtime_dir() / "balance-watch.json"
|
||||||
|
|
||||||
|
|
||||||
|
def utc_now_iso() -> str:
|
||||||
|
return datetime.now(timezone.utc).replace(microsecond=0).isoformat()
|
||||||
|
|
||||||
|
|
||||||
|
def load_balance_state() -> BalanceWatchState | None:
|
||||||
|
path = balance_watch_path()
|
||||||
|
if not path.is_file():
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
raw = json.loads(path.read_text(encoding="utf-8"))
|
||||||
|
except (OSError, json.JSONDecodeError):
|
||||||
|
return None
|
||||||
|
if not isinstance(raw, dict) or "baseline_rub" not in raw:
|
||||||
|
return None
|
||||||
|
return BalanceWatchState.from_dict(raw)
|
||||||
|
|
||||||
|
|
||||||
|
def save_balance_state(state: BalanceWatchState) -> None:
|
||||||
|
runtime_dir().mkdir(parents=True, exist_ok=True)
|
||||||
|
state.updated_at = utc_now_iso()
|
||||||
|
balance_watch_path().write_text(
|
||||||
|
json.dumps(state.to_dict(), indent=2) + "\n", encoding="utf-8"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def clear_balance_state() -> None:
|
||||||
|
path = balance_watch_path()
|
||||||
|
if path.is_file():
|
||||||
|
path.unlink()
|
||||||
|
|
||||||
|
|
||||||
|
def balance_rub_from_payload(payload: dict[str, Any]) -> float:
|
||||||
|
"""Selectel returns money as integer kopecks → rubles."""
|
||||||
|
data = payload.get("data") if isinstance(payload.get("data"), dict) else payload
|
||||||
|
billings = data.get("billings") if isinstance(data, dict) else None
|
||||||
|
if not isinstance(billings, list) or not billings:
|
||||||
|
raise CloudError("Selectel balances: нет data.billings")
|
||||||
|
total_kop = 0
|
||||||
|
for row in billings:
|
||||||
|
if not isinstance(row, dict):
|
||||||
|
continue
|
||||||
|
if row.get("balances_values_sum") is not None:
|
||||||
|
total_kop += int(row["balances_values_sum"])
|
||||||
|
elif row.get("final_sum") is not None:
|
||||||
|
total_kop += int(row["final_sum"])
|
||||||
|
return total_kop / 100.0
|
||||||
|
|
||||||
|
|
||||||
|
def fetch_balance_rub(api_token: str, *, timeout: float = 20.0) -> float:
|
||||||
|
token = (api_token or "").strip()
|
||||||
|
if not token:
|
||||||
|
raise CloudError("SELECTEL_API_TOKEN пуст — баланс недоступен")
|
||||||
|
headers = {"X-Token": token, "Accept": "application/json"}
|
||||||
|
try:
|
||||||
|
with httpx.Client(timeout=timeout, follow_redirects=True) as client:
|
||||||
|
resp = client.get(BALANCES_URL, headers=headers)
|
||||||
|
except httpx.HTTPError as exc:
|
||||||
|
raise CloudError(f"Selectel balances сеть: {exc}") from exc
|
||||||
|
if resp.status_code in {401, 403}:
|
||||||
|
raise CloudError(
|
||||||
|
f"Selectel balances HTTP {resp.status_code} — нужен статический "
|
||||||
|
"API-ключ аккаунта (X-Token), не OpenStack пароль"
|
||||||
|
)
|
||||||
|
if resp.status_code != 200:
|
||||||
|
raise CloudError(f"Selectel balances HTTP {resp.status_code}: {resp.text[:200]}")
|
||||||
|
try:
|
||||||
|
payload = resp.json()
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise CloudError("Selectel balances: не JSON") from exc
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise CloudError("Selectel balances: неожиданный JSON")
|
||||||
|
return balance_rub_from_payload(payload)
|
||||||
|
|
||||||
|
|
||||||
|
def spend_steps(baseline_rub: float, current_rub: float, step_rub: float) -> int:
|
||||||
|
"""How many full step intervals have been spent since baseline."""
|
||||||
|
if step_rub <= 0:
|
||||||
|
return 0
|
||||||
|
spend = baseline_rub - current_rub
|
||||||
|
if spend < 0:
|
||||||
|
return 0
|
||||||
|
return int(spend // step_rub)
|
||||||
|
|
||||||
|
|
||||||
|
def plan_balance_notices(
|
||||||
|
state: BalanceWatchState,
|
||||||
|
current_rub: float,
|
||||||
|
*,
|
||||||
|
step_rub: float = DEFAULT_STEP_RUB,
|
||||||
|
low_rub: float = 0.0,
|
||||||
|
) -> tuple[BalanceWatchState, list[str]]:
|
||||||
|
"""
|
||||||
|
Update state for top-ups / spend steps / low watermark.
|
||||||
|
Returns (new_state, human messages to show).
|
||||||
|
"""
|
||||||
|
messages: list[str] = []
|
||||||
|
# Top-up: reset baseline so steps restart from new balance.
|
||||||
|
if current_rub > state.baseline_rub + 0.01:
|
||||||
|
state = BalanceWatchState(
|
||||||
|
baseline_rub=current_rub,
|
||||||
|
last_notified_step=0,
|
||||||
|
low_notified=False,
|
||||||
|
last_balance_rub=current_rub,
|
||||||
|
)
|
||||||
|
messages.append(f"Баланс пополнен: {current_rub:.0f} ₽ (новый baseline)")
|
||||||
|
return state, messages
|
||||||
|
|
||||||
|
crossed = spend_steps(state.baseline_rub, current_rub, step_rub)
|
||||||
|
while state.last_notified_step < crossed:
|
||||||
|
state.last_notified_step += 1
|
||||||
|
spent = state.last_notified_step * step_rub
|
||||||
|
messages.append(
|
||||||
|
f"Списано ~{spent:.0f} ₽ с момента up "
|
||||||
|
f"(осталось {current_rub:.0f} ₽, шаг {step_rub:.0f} ₽)"
|
||||||
|
)
|
||||||
|
|
||||||
|
if low_rub > 0 and current_rub < low_rub and not state.low_notified:
|
||||||
|
state.low_notified = True
|
||||||
|
messages.append(f"Баланс низкий: {current_rub:.0f} ₽ (< {low_rub:.0f} ₽)")
|
||||||
|
|
||||||
|
state.last_balance_rub = current_rub
|
||||||
|
return state, messages
|
||||||
@@ -51,6 +51,12 @@ def _as_int(value: str | None, default: int) -> int:
|
|||||||
return int(value)
|
return int(value)
|
||||||
|
|
||||||
|
|
||||||
|
def _as_float(value: str | None, default: float) -> float:
|
||||||
|
if value is None or str(value).strip() == "":
|
||||||
|
return default
|
||||||
|
return float(str(value).strip().replace(",", "."))
|
||||||
|
|
||||||
|
|
||||||
def _csv(value: str | None, default: tuple[str, ...]) -> tuple[str, ...]:
|
def _csv(value: str | None, default: tuple[str, ...]) -> tuple[str, ...]:
|
||||||
if value is None or value.strip() == "":
|
if value is None or value.strip() == "":
|
||||||
return default
|
return default
|
||||||
@@ -114,6 +120,9 @@ class Config:
|
|||||||
idle_grace_minutes: int
|
idle_grace_minutes: int
|
||||||
pull_output: bool
|
pull_output: bool
|
||||||
notify_ready: bool
|
notify_ready: bool
|
||||||
|
selectel_api_token: str
|
||||||
|
balance_notify_step_rub: float
|
||||||
|
balance_notify_low_rub: float
|
||||||
|
|
||||||
missing: list[str] = field(default_factory=list)
|
missing: list[str] = field(default_factory=list)
|
||||||
|
|
||||||
@@ -245,5 +254,8 @@ def load_config(*, require_auth: bool = True) -> Config:
|
|||||||
idle_grace_minutes=_as_int(os.environ.get("IDLE_GRACE_MINUTES"), 45),
|
idle_grace_minutes=_as_int(os.environ.get("IDLE_GRACE_MINUTES"), 45),
|
||||||
pull_output=_as_bool(os.environ.get("PULL_OUTPUT"), False),
|
pull_output=_as_bool(os.environ.get("PULL_OUTPUT"), False),
|
||||||
notify_ready=_as_bool(os.environ.get("NOTIFY_READY"), True),
|
notify_ready=_as_bool(os.environ.get("NOTIFY_READY"), True),
|
||||||
|
selectel_api_token=(os.environ.get("SELECTEL_API_TOKEN") or "").strip(),
|
||||||
|
balance_notify_step_rub=_as_float(os.environ.get("BALANCE_NOTIFY_STEP_RUB"), 200.0),
|
||||||
|
balance_notify_low_rub=_as_float(os.environ.get("BALANCE_NOTIFY_LOW_RUB"), 0.0),
|
||||||
missing=missing,
|
missing=missing,
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -272,6 +272,7 @@ def install_watchdog(
|
|||||||
f"local-watchdog установлен ({platform}): тик каждые {interval} мин. "
|
f"local-watchdog установлен ({platform}): тик каждые {interval} мин. "
|
||||||
f"Grace {grace_seconds() // 60} мин после смерти процесса туннеля → stop. "
|
f"Grace {grace_seconds() // 60} мин после смерти процесса туннеля → stop. "
|
||||||
f"Ctrl+C на туннеле GPU не гасит. "
|
f"Ctrl+C на туннеле GPU не гасит. "
|
||||||
|
"При SELECTEL_API_TOKEN — toast каждые BALANCE_NOTIFY_STEP_RUB ₽ (дефолт 200)."
|
||||||
)
|
)
|
||||||
return marker
|
return marker
|
||||||
|
|
||||||
@@ -314,9 +315,64 @@ def watchdog_status_lines() -> list[str]:
|
|||||||
f"pid={lease.pid} alive={alive} hb={lease.heartbeat_at}"
|
f"pid={lease.pid} alive={alive} hb={lease.heartbeat_at}"
|
||||||
)
|
)
|
||||||
lines.append(f"grace: {grace_seconds() // 60} мин (LOCAL_WATCHDOG_GRACE_MINUTES)")
|
lines.append(f"grace: {grace_seconds() // 60} мин (LOCAL_WATCHDOG_GRACE_MINUTES)")
|
||||||
|
try:
|
||||||
|
from gpu_rent.balance import load_balance_state
|
||||||
|
|
||||||
|
bal = load_balance_state()
|
||||||
|
if bal:
|
||||||
|
lines.append(
|
||||||
|
f"balance: baseline {bal.baseline_rub:.0f} ₽ "
|
||||||
|
f"last={bal.last_balance_rub} step_notified={bal.last_notified_step}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
lines.append("balance: нет baseline (появится на up при SELECTEL_API_TOKEN)")
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
return lines
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def check_balance_on_tick(cfg: Config, *, log: Log = print) -> list[str]:
|
||||||
|
"""Fetch balance; toast on each BALANCE_NOTIFY_STEP_RUB spent since baseline.
|
||||||
|
|
||||||
|
No-op without SELECTEL_API_TOKEN. Arms baseline on first successful fetch
|
||||||
|
if missing but a GPU session may be running.
|
||||||
|
"""
|
||||||
|
from gpu_rent.balance import (
|
||||||
|
BalanceWatchState,
|
||||||
|
fetch_balance_rub,
|
||||||
|
load_balance_state,
|
||||||
|
plan_balance_notices,
|
||||||
|
save_balance_state,
|
||||||
|
)
|
||||||
|
from gpu_rent.notify import notify_message
|
||||||
|
|
||||||
|
token = (cfg.selectel_api_token or "").strip()
|
||||||
|
if not token:
|
||||||
|
return []
|
||||||
|
step = float(cfg.balance_notify_step_rub or 200.0)
|
||||||
|
low = float(cfg.balance_notify_low_rub or 0.0)
|
||||||
|
try:
|
||||||
|
current = fetch_balance_rub(token)
|
||||||
|
except Exception as exc:
|
||||||
|
log(f"balance: {exc}")
|
||||||
|
return []
|
||||||
|
|
||||||
|
state = load_balance_state()
|
||||||
|
if state is None:
|
||||||
|
state = BalanceWatchState(baseline_rub=current, last_balance_rub=current)
|
||||||
|
save_balance_state(state)
|
||||||
|
log(f"balance: baseline {current:.0f} ₽ (шаг {step:.0f} ₽)")
|
||||||
|
return []
|
||||||
|
|
||||||
|
state, messages = plan_balance_notices(
|
||||||
|
state, current, step_rub=step, low_rub=low
|
||||||
|
)
|
||||||
|
save_balance_state(state)
|
||||||
|
for msg in messages:
|
||||||
|
notify_message("gpu-rent баланс", msg, log)
|
||||||
|
return messages
|
||||||
|
|
||||||
|
|
||||||
def run_tick(*, dry_run: bool = False, log: Log = print) -> TickDecision:
|
def run_tick(*, dry_run: bool = False, log: Log = print) -> TickDecision:
|
||||||
from gpu_rent.session import cmd_stop
|
from gpu_rent.session import cmd_stop
|
||||||
from gpu_rent.state import load_state
|
from gpu_rent.state import load_state
|
||||||
@@ -331,6 +387,18 @@ def run_tick(*, dry_run: bool = False, log: Log = print) -> TickDecision:
|
|||||||
process_alive=pid_alive(lease.pid if lease else None),
|
process_alive=pid_alive(lease.pid if lease else None),
|
||||||
grace_sec=grace_seconds(),
|
grace_sec=grace_seconds(),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# Balance steps run on every tick while watchdog is installed (same timer).
|
||||||
|
if watchdog_installed() and not dry_run:
|
||||||
|
try:
|
||||||
|
from gpu_rent.config import load_config
|
||||||
|
|
||||||
|
cfg = load_config(require_auth=False)
|
||||||
|
if state.server_id or (lease and lease.armed):
|
||||||
|
check_balance_on_tick(cfg, log=log)
|
||||||
|
except Exception as exc:
|
||||||
|
log(f"balance tick: {exc}")
|
||||||
|
|
||||||
if decision.kind == "noop":
|
if decision.kind == "noop":
|
||||||
log(f"watchdog tick: noop ({decision.detail})")
|
log(f"watchdog tick: noop ({decision.detail})")
|
||||||
return decision
|
return decision
|
||||||
|
|||||||
+21
-6
@@ -1,4 +1,4 @@
|
|||||||
"""Ready notification: toast + sound. Never raises."""
|
"""Ready / balance notifications: toast + sound. Never raises."""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
@@ -10,7 +10,12 @@ from gpu_rent.config import Config
|
|||||||
|
|
||||||
Log = Callable[[str], None]
|
Log = Callable[[str], None]
|
||||||
|
|
||||||
__all__ = ["notify_ready", "print_mcp_snippet", "print_access_card"]
|
__all__ = [
|
||||||
|
"notify_ready",
|
||||||
|
"notify_message",
|
||||||
|
"print_mcp_snippet",
|
||||||
|
"print_access_card",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
def notify_ready(cfg: Config, log: Log) -> None:
|
def notify_ready(cfg: Config, log: Log) -> None:
|
||||||
@@ -19,7 +24,15 @@ def notify_ready(cfg: Config, log: Log) -> None:
|
|||||||
log("NOTIFY_READY: SwarmUI Idle")
|
log("NOTIFY_READY: SwarmUI Idle")
|
||||||
_sound()
|
_sound()
|
||||||
if sys.platform == "win32":
|
if sys.platform == "win32":
|
||||||
_windows_toast(log)
|
_windows_toast("SwarmUI backend Idle — gpu-rent tunnel", log)
|
||||||
|
|
||||||
|
|
||||||
|
def notify_message(title: str, body: str, log: Log) -> None:
|
||||||
|
"""Generic toast/sound for balance steps etc."""
|
||||||
|
log(f"{title}: {body}")
|
||||||
|
_sound()
|
||||||
|
if sys.platform == "win32":
|
||||||
|
_windows_toast(body, log, title=title)
|
||||||
|
|
||||||
|
|
||||||
def _sound() -> None:
|
def _sound() -> None:
|
||||||
@@ -38,15 +51,17 @@ def _sound() -> None:
|
|||||||
pass
|
pass
|
||||||
|
|
||||||
|
|
||||||
def _windows_toast(log: Log) -> None:
|
def _windows_toast(body: str, log: Log, *, title: str = "gpu-rent") -> None:
|
||||||
|
safe_title = title.replace("'", "''")
|
||||||
|
safe_body = body.replace("'", "''")
|
||||||
script = (
|
script = (
|
||||||
"[Windows.UI.Notifications.ToastNotificationManager, Windows.UI.Notifications, "
|
"[Windows.UI.Notifications.ToastNotificationManager, Windows.UI.Notifications, "
|
||||||
"ContentType = WindowsRuntime] > $null; "
|
"ContentType = WindowsRuntime] > $null; "
|
||||||
"$template = [Windows.UI.Notifications.ToastNotificationManager]::GetTemplateContent("
|
"$template = [Windows.UI.Notifications.ToastNotificationManager]::GetTemplateContent("
|
||||||
"[Windows.UI.Notifications.ToastTemplateType]::ToastText02); "
|
"[Windows.UI.Notifications.ToastTemplateType]::ToastText02); "
|
||||||
"$text = $template.GetElementsByTagName('text'); "
|
"$text = $template.GetElementsByTagName('text'); "
|
||||||
"$text.Item(0).AppendChild($template.CreateTextNode('gpu-rent')) | Out-Null; "
|
f"$text.Item(0).AppendChild($template.CreateTextNode('{safe_title}')) | Out-Null; "
|
||||||
"$text.Item(1).AppendChild($template.CreateTextNode('SwarmUI backend Idle — gpu-rent tunnel')) | Out-Null; "
|
f"$text.Item(1).AppendChild($template.CreateTextNode('{safe_body}')) | Out-Null; "
|
||||||
"$toast = [Windows.UI.Notifications.ToastNotification]::new($template); "
|
"$toast = [Windows.UI.Notifications.ToastNotification]::new($template); "
|
||||||
"[Windows.UI.Notifications.ToastNotificationManager]::CreateToastNotifier('gpu-rent').Show($toast)"
|
"[Windows.UI.Notifications.ToastNotificationManager]::CreateToastNotifier('gpu-rent').Show($toast)"
|
||||||
)
|
)
|
||||||
|
|||||||
+295
-1
@@ -1,12 +1,21 @@
|
|||||||
"""Wait until SwarmUI HTTP is up and backend is Idle (on the VM)."""
|
"""Wait until SwarmUI HTTP is up and backend is Idle (on the VM).
|
||||||
|
|
||||||
|
Also end-of-up stack verification: every enabled service must answer.
|
||||||
|
"""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import socket
|
||||||
import time
|
import time
|
||||||
|
import urllib.error
|
||||||
|
import urllib.request
|
||||||
from collections.abc import Callable
|
from collections.abc import Callable
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
from gpu_rent.config import Config
|
from gpu_rent.config import Config
|
||||||
from gpu_rent.errors import CloudError
|
from gpu_rent.errors import CloudError
|
||||||
|
from gpu_rent.llm_runtime import normalize_runtime
|
||||||
from gpu_rent.ssh_ops import run_ssh
|
from gpu_rent.ssh_ops import run_ssh
|
||||||
|
|
||||||
Log = Callable[[str], None]
|
Log = Callable[[str], None]
|
||||||
@@ -57,6 +66,93 @@ while time.time() < deadline:
|
|||||||
print("WAIT timeout-slice")
|
print("WAIT timeout-slice")
|
||||||
"""
|
"""
|
||||||
|
|
||||||
|
# One-shot probe of configured stack endpoints on the VM (JSON line).
|
||||||
|
_REMOTE_STACK_PROBE = r'''
|
||||||
|
import json, urllib.error, urllib.request, subprocess
|
||||||
|
|
||||||
|
def http_ok(url, timeout=4.0):
|
||||||
|
try:
|
||||||
|
req = urllib.request.Request(url, method="GET")
|
||||||
|
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||||
|
code = getattr(resp, "status", 200) or 200
|
||||||
|
body = resp.read(256)
|
||||||
|
return True, f"HTTP {code} ({len(body)}b)"
|
||||||
|
except Exception as exc:
|
||||||
|
return False, str(exc)[:160]
|
||||||
|
|
||||||
|
def unit_active(name):
|
||||||
|
try:
|
||||||
|
out = subprocess.check_output(
|
||||||
|
["systemctl", "is-active", name],
|
||||||
|
text=True,
|
||||||
|
stderr=subprocess.DEVNULL,
|
||||||
|
).strip()
|
||||||
|
return out
|
||||||
|
except Exception:
|
||||||
|
return "unknown"
|
||||||
|
|
||||||
|
checks = []
|
||||||
|
want_swarm = WANT_SWARM
|
||||||
|
want_ollama = WANT_OLLAMA
|
||||||
|
want_llama = WANT_LLAMA
|
||||||
|
|
||||||
|
if want_swarm:
|
||||||
|
ok, detail = http_ok("http://127.0.0.1:7801/")
|
||||||
|
if not ok:
|
||||||
|
# API may answer when / does not
|
||||||
|
ok2, d2 = False, detail
|
||||||
|
try:
|
||||||
|
req = urllib.request.Request(
|
||||||
|
"http://127.0.0.1:7801/API/GetNewSession",
|
||||||
|
data=b"{}",
|
||||||
|
headers={"Content-Type": "application/json"},
|
||||||
|
method="POST",
|
||||||
|
)
|
||||||
|
with urllib.request.urlopen(req, timeout=5) as resp:
|
||||||
|
ok2 = True
|
||||||
|
d2 = f"API session HTTP {getattr(resp, 'status', 200)}"
|
||||||
|
except Exception as exc:
|
||||||
|
d2 = str(exc)[:160]
|
||||||
|
ok, detail = ok2, d2
|
||||||
|
checks.append({
|
||||||
|
"name": "swarmui",
|
||||||
|
"ok": ok,
|
||||||
|
"detail": detail,
|
||||||
|
"unit": unit_active("swarmui"),
|
||||||
|
})
|
||||||
|
|
||||||
|
if want_ollama:
|
||||||
|
ok, detail = http_ok("http://127.0.0.1:11434/api/tags")
|
||||||
|
checks.append({
|
||||||
|
"name": "ollama",
|
||||||
|
"ok": ok,
|
||||||
|
"detail": detail,
|
||||||
|
"unit": unit_active("gpu-rent-ollama"),
|
||||||
|
})
|
||||||
|
|
||||||
|
if want_llama:
|
||||||
|
ok, detail = http_ok("http://127.0.0.1:8080/health")
|
||||||
|
if not ok:
|
||||||
|
ok2, d2 = http_ok("http://127.0.0.1:8080/v1/models")
|
||||||
|
ok, detail = ok2, d2
|
||||||
|
checks.append({
|
||||||
|
"name": "llamacpp",
|
||||||
|
"ok": ok,
|
||||||
|
"detail": detail,
|
||||||
|
"unit": unit_active("gpu-rent-llamacpp"),
|
||||||
|
})
|
||||||
|
|
||||||
|
print(json.dumps({"checks": checks}, ensure_ascii=False))
|
||||||
|
'''
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ServiceCheck:
|
||||||
|
name: str
|
||||||
|
ok: bool
|
||||||
|
detail: str
|
||||||
|
where: str = "vm" # vm | local
|
||||||
|
|
||||||
|
|
||||||
def wait_backend_idle(
|
def wait_backend_idle(
|
||||||
cfg: Config,
|
cfg: Config,
|
||||||
@@ -93,3 +189,201 @@ def wait_backend_idle(
|
|||||||
f"backend не стал Idle за {int(timeout)} с. "
|
f"backend не стал Idle за {int(timeout)} с. "
|
||||||
"Проверь journalctl -u swarmui на VM; GPU всё ещё жив."
|
"Проверь journalctl -u swarmui на VM; GPU всё ещё жив."
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _expected_services(cfg: Config) -> tuple[bool, bool, bool]:
|
||||||
|
swarm = bool(getattr(cfg, "enable_swarmui", True))
|
||||||
|
rt = normalize_runtime(getattr(cfg, "llm_runtime", "none"))
|
||||||
|
return swarm, rt == "ollama", rt == "llamacpp"
|
||||||
|
|
||||||
|
|
||||||
|
def _probe_vm_once(cfg: Config, host: str) -> list[ServiceCheck]:
|
||||||
|
want_swarm, want_ollama, want_llama = _expected_services(cfg)
|
||||||
|
script = (
|
||||||
|
_REMOTE_STACK_PROBE.replace("WANT_SWARM", "True" if want_swarm else "False")
|
||||||
|
.replace("WANT_OLLAMA", "True" if want_ollama else "False")
|
||||||
|
.replace("WANT_LLAMA", "True" if want_llama else "False")
|
||||||
|
)
|
||||||
|
out = run_ssh(
|
||||||
|
cfg,
|
||||||
|
host,
|
||||||
|
"python3 - <<'PY'\n" + script + "\nPY",
|
||||||
|
check=False,
|
||||||
|
timeout=60,
|
||||||
|
).strip()
|
||||||
|
line = ""
|
||||||
|
for row in reversed(out.splitlines()):
|
||||||
|
row = row.strip()
|
||||||
|
if row.startswith("{"):
|
||||||
|
line = row
|
||||||
|
break
|
||||||
|
if not line:
|
||||||
|
return [
|
||||||
|
ServiceCheck("stack", False, f"нет JSON от probe: {out[-200:]}", "vm")
|
||||||
|
]
|
||||||
|
try:
|
||||||
|
data = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
return [ServiceCheck("stack", False, f"битый JSON: {line[:200]}", "vm")]
|
||||||
|
checks: list[ServiceCheck] = []
|
||||||
|
for item in data.get("checks") or []:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
name = str(item.get("name") or "?")
|
||||||
|
detail = str(item.get("detail") or "")
|
||||||
|
unit = str(item.get("unit") or "")
|
||||||
|
if unit and unit != "unknown":
|
||||||
|
detail = f"{detail}; unit={unit}"
|
||||||
|
checks.append(
|
||||||
|
ServiceCheck(name=name, ok=bool(item.get("ok")), detail=detail, where="vm")
|
||||||
|
)
|
||||||
|
return checks
|
||||||
|
|
||||||
|
|
||||||
|
def verify_stack_on_vm(
|
||||||
|
cfg: Config,
|
||||||
|
host: str,
|
||||||
|
log: Log,
|
||||||
|
*,
|
||||||
|
timeout: float = 300.0,
|
||||||
|
poll_every: float = 8.0,
|
||||||
|
raise_on_fail: bool = True,
|
||||||
|
) -> list[ServiceCheck]:
|
||||||
|
"""Poll until every enabled service answers on the VM loopback."""
|
||||||
|
want_swarm, want_ollama, want_llama = _expected_services(cfg)
|
||||||
|
if not (want_swarm or want_ollama or want_llama):
|
||||||
|
log("проверка стека: нечего ждать (swarm off, LLM none)")
|
||||||
|
return []
|
||||||
|
|
||||||
|
names = []
|
||||||
|
if want_swarm:
|
||||||
|
names.append("SwarmUI :7801")
|
||||||
|
if want_ollama:
|
||||||
|
names.append("Ollama :11434")
|
||||||
|
if want_llama:
|
||||||
|
names.append("llama.cpp :8080")
|
||||||
|
log(f"проверка на VM: {', '.join(names)}")
|
||||||
|
|
||||||
|
deadline = time.time() + timeout
|
||||||
|
last: list[ServiceCheck] = []
|
||||||
|
while time.time() < deadline:
|
||||||
|
try:
|
||||||
|
last = _probe_vm_once(cfg, host)
|
||||||
|
except Exception as exc:
|
||||||
|
last = [ServiceCheck("ssh", False, str(exc)[:200], "vm")]
|
||||||
|
if last and all(c.ok for c in last):
|
||||||
|
for c in last:
|
||||||
|
log(f" [ok] {c.name}: {c.detail}")
|
||||||
|
log("проверка VM: всё отвечает")
|
||||||
|
return last
|
||||||
|
bad = ", ".join(f"{c.name}={c.detail}" for c in last if not c.ok) or "пусто"
|
||||||
|
log(f" … ещё нет: {bad}")
|
||||||
|
time.sleep(poll_every)
|
||||||
|
|
||||||
|
for c in last:
|
||||||
|
mark = "ok" if c.ok else "FAIL"
|
||||||
|
log(f" [{mark}] {c.name}: {c.detail}")
|
||||||
|
if raise_on_fail and any(not c.ok for c in last):
|
||||||
|
failed = [c.name for c in last if not c.ok]
|
||||||
|
raise CloudError(
|
||||||
|
f"сервисы не ответили на VM за {int(timeout)} с: {', '.join(failed)}. "
|
||||||
|
"GPU жив — смотри journalctl / gpu-rent logs"
|
||||||
|
)
|
||||||
|
return last
|
||||||
|
|
||||||
|
|
||||||
|
def _tcp_ok(port: int, host: str = "127.0.0.1", timeout: float = 0.8) -> bool:
|
||||||
|
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
||||||
|
sock.settimeout(timeout)
|
||||||
|
try:
|
||||||
|
return sock.connect_ex((host, port)) == 0
|
||||||
|
finally:
|
||||||
|
sock.close()
|
||||||
|
|
||||||
|
|
||||||
|
def _http_local(url: str, timeout: float = 3.0) -> tuple[bool, str]:
|
||||||
|
try:
|
||||||
|
req = urllib.request.Request(url, method="GET")
|
||||||
|
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||||
|
code = getattr(resp, "status", 200) or 200
|
||||||
|
return True, f"HTTP {code}"
|
||||||
|
except Exception as exc:
|
||||||
|
return False, str(exc)[:160]
|
||||||
|
|
||||||
|
|
||||||
|
def verify_stack_local(
|
||||||
|
cfg: Config,
|
||||||
|
log: Log,
|
||||||
|
*,
|
||||||
|
timeout: float = 60.0,
|
||||||
|
poll_every: float = 2.0,
|
||||||
|
raise_on_fail: bool = True,
|
||||||
|
) -> list[ServiceCheck]:
|
||||||
|
"""After tunnel: local ports + light HTTP for enabled services."""
|
||||||
|
want_swarm, want_ollama, want_llama = _expected_services(cfg)
|
||||||
|
targets: list[tuple[str, int, str | None]] = []
|
||||||
|
if want_swarm:
|
||||||
|
targets.append(("swarmui", int(cfg.swarmui_local_port), None))
|
||||||
|
if want_ollama:
|
||||||
|
targets.append(
|
||||||
|
("ollama", int(cfg.ollama_local_port), f"http://127.0.0.1:{cfg.ollama_local_port}/api/tags")
|
||||||
|
)
|
||||||
|
if want_llama:
|
||||||
|
p = int(cfg.llamacpp_local_port)
|
||||||
|
targets.append(("llamacpp", p, f"http://127.0.0.1:{p}/health"))
|
||||||
|
|
||||||
|
if not targets:
|
||||||
|
return []
|
||||||
|
|
||||||
|
log("проверка туннеля (localhost): " + ", ".join(f"{n}:{port}" for n, port, _ in targets))
|
||||||
|
deadline = time.time() + timeout
|
||||||
|
last: list[ServiceCheck] = []
|
||||||
|
while time.time() < deadline:
|
||||||
|
last = []
|
||||||
|
for name, port, url in targets:
|
||||||
|
if not _tcp_ok(port):
|
||||||
|
last.append(ServiceCheck(name, False, f"порт {port} закрыт", "local"))
|
||||||
|
continue
|
||||||
|
if url:
|
||||||
|
ok, detail = _http_local(url)
|
||||||
|
if not ok and name == "llamacpp":
|
||||||
|
ok, detail = _http_local(
|
||||||
|
f"http://127.0.0.1:{port}/v1/models"
|
||||||
|
)
|
||||||
|
last.append(ServiceCheck(name, ok, detail, "local"))
|
||||||
|
else:
|
||||||
|
# SwarmUI: TCP enough (API may need POST); try API too
|
||||||
|
ok, detail = _http_local(f"http://127.0.0.1:{port}/")
|
||||||
|
if not ok:
|
||||||
|
try:
|
||||||
|
req = urllib.request.Request(
|
||||||
|
f"http://127.0.0.1:{port}/API/GetNewSession",
|
||||||
|
data=b"{}",
|
||||||
|
headers={"Content-Type": "application/json"},
|
||||||
|
method="POST",
|
||||||
|
)
|
||||||
|
with urllib.request.urlopen(req, timeout=4) as resp:
|
||||||
|
ok = True
|
||||||
|
detail = f"API HTTP {getattr(resp, 'status', 200)}"
|
||||||
|
except Exception as exc:
|
||||||
|
detail = f"TCP ok; HTTP {exc}"[:160]
|
||||||
|
# TCP open is enough for swarm local check
|
||||||
|
ok = True
|
||||||
|
detail = f"TCP :{port} open"
|
||||||
|
last.append(ServiceCheck(name, ok, detail, "local"))
|
||||||
|
if last and all(c.ok for c in last):
|
||||||
|
for c in last:
|
||||||
|
log(f" [ok] localhost {c.name}: {c.detail}")
|
||||||
|
log("проверка туннеля: всё доступно")
|
||||||
|
return last
|
||||||
|
time.sleep(poll_every)
|
||||||
|
|
||||||
|
for c in last:
|
||||||
|
mark = "ok" if c.ok else "FAIL"
|
||||||
|
log(f" [{mark}] localhost {c.name}: {c.detail}")
|
||||||
|
if raise_on_fail and any(not c.ok for c in last):
|
||||||
|
failed = [c.name for c in last if not c.ok]
|
||||||
|
raise CloudError(
|
||||||
|
f"туннель поднят, но локально не отвечает: {', '.join(failed)}"
|
||||||
|
)
|
||||||
|
return last
|
||||||
|
|||||||
+41
-1
@@ -22,7 +22,7 @@ from gpu_rent.cloud import (
|
|||||||
)
|
)
|
||||||
from gpu_rent.bootstrap import run_bootstrap
|
from gpu_rent.bootstrap import run_bootstrap
|
||||||
from gpu_rent.provision import provision_vm, tune_swarm_perf
|
from gpu_rent.provision import provision_vm, tune_swarm_perf
|
||||||
from gpu_rent.ready import wait_backend_idle
|
from gpu_rent.ready import verify_stack_on_vm, wait_backend_idle
|
||||||
from gpu_rent.snapshot import ensure_boot_snapshot
|
from gpu_rent.snapshot import ensure_boot_snapshot
|
||||||
from gpu_rent.notify import notify_ready
|
from gpu_rent.notify import notify_ready
|
||||||
from gpu_rent.config import Config
|
from gpu_rent.config import Config
|
||||||
@@ -127,6 +127,19 @@ def _bind_access(
|
|||||||
log(f"perf tune: {exc}")
|
log(f"perf tune: {exc}")
|
||||||
else:
|
else:
|
||||||
log("ready: llm-only (без ожидания SwarmUI Idle)")
|
log("ready: llm-only (без ожидания SwarmUI Idle)")
|
||||||
|
|
||||||
|
try:
|
||||||
|
checks = verify_stack_on_vm(cfg, ip, log, timeout=300.0)
|
||||||
|
state.notes = dict(state.notes or {})
|
||||||
|
state.notes["stack_vm"] = [
|
||||||
|
{"name": c.name, "ok": c.ok, "detail": c.detail} for c in checks
|
||||||
|
]
|
||||||
|
except CloudError as exc:
|
||||||
|
state.notes = dict(state.notes or {})
|
||||||
|
state.notes["stack_vm_error"] = str(exc)[:500]
|
||||||
|
save_state(state)
|
||||||
|
raise
|
||||||
|
|
||||||
try:
|
try:
|
||||||
ensure_boot_snapshot(
|
ensure_boot_snapshot(
|
||||||
conn,
|
conn,
|
||||||
@@ -137,6 +150,27 @@ def _bind_access(
|
|||||||
except CloudError as exc:
|
except CloudError as exc:
|
||||||
log(f"snapshot: {exc}")
|
log(f"snapshot: {exc}")
|
||||||
notify_ready(cfg, log)
|
notify_ready(cfg, log)
|
||||||
|
try:
|
||||||
|
from gpu_rent.balance import (
|
||||||
|
BalanceWatchState,
|
||||||
|
fetch_balance_rub,
|
||||||
|
save_balance_state,
|
||||||
|
)
|
||||||
|
|
||||||
|
if cfg.selectel_api_token:
|
||||||
|
rub = fetch_balance_rub(cfg.selectel_api_token)
|
||||||
|
save_balance_state(
|
||||||
|
BalanceWatchState(baseline_rub=rub, last_balance_rub=rub)
|
||||||
|
)
|
||||||
|
log(
|
||||||
|
f"balance: baseline {rub:.0f} ₽ "
|
||||||
|
f"(уведомление каждые {cfg.balance_notify_step_rub:.0f} ₽, "
|
||||||
|
"нужен local-watchdog)"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
log("balance: SELECTEL_API_TOKEN не задан — шаги по ₽ skip")
|
||||||
|
except Exception as exc:
|
||||||
|
log(f"balance baseline: {exc}")
|
||||||
state.bootstrapped = True
|
state.bootstrapped = True
|
||||||
state.phase = "ready_cloud"
|
state.phase = "ready_cloud"
|
||||||
state.notes = dict(state.notes or {})
|
state.notes = dict(state.notes or {})
|
||||||
@@ -464,5 +498,11 @@ def cmd_stop(
|
|||||||
clear_lease()
|
clear_lease()
|
||||||
except Exception:
|
except Exception:
|
||||||
pass
|
pass
|
||||||
|
try:
|
||||||
|
from gpu_rent.balance import clear_balance_state
|
||||||
|
|
||||||
|
clear_balance_state()
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
log("фаза idle" + ("" if destroy_disks else " (диски на месте)"))
|
log("фаза idle" + ("" if destroy_disks else " (диски на месте)"))
|
||||||
return state
|
return state
|
||||||
|
|||||||
@@ -194,6 +194,20 @@ def run_tunnel(
|
|||||||
else:
|
else:
|
||||||
open_url = f"http://127.0.0.1:{cfg.swarmui_local_port}"
|
open_url = f"http://127.0.0.1:{cfg.swarmui_local_port}"
|
||||||
|
|
||||||
|
from gpu_rent.ready import verify_stack_local
|
||||||
|
|
||||||
|
try:
|
||||||
|
local_checks = verify_stack_local(cfg, log, timeout=90.0)
|
||||||
|
state = load_state()
|
||||||
|
state.notes = dict(state.notes or {})
|
||||||
|
state.notes["stack_local"] = [
|
||||||
|
{"name": c.name, "ok": c.ok, "detail": c.detail} for c in local_checks
|
||||||
|
]
|
||||||
|
save_state(state)
|
||||||
|
except CloudError as exc:
|
||||||
|
log(f"проверка туннеля: {exc}")
|
||||||
|
raise
|
||||||
|
|
||||||
from gpu_rent.access_card import print_access_card
|
from gpu_rent.access_card import print_access_card
|
||||||
|
|
||||||
print_access_card(cfg, tunneled=True, host=current_host)
|
print_access_card(cfg, tunneled=True, host=current_host)
|
||||||
|
|||||||
@@ -0,0 +1,56 @@
|
|||||||
|
from gpu_rent.balance import (
|
||||||
|
BalanceWatchState,
|
||||||
|
balance_rub_from_payload,
|
||||||
|
plan_balance_notices,
|
||||||
|
spend_steps,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_balance_kopecks():
|
||||||
|
payload = {
|
||||||
|
"data": {
|
||||||
|
"billings": [
|
||||||
|
{"balances_values_sum": 123456}, # 1234.56 ₽
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
assert balance_rub_from_payload(payload) == 1234.56
|
||||||
|
|
||||||
|
|
||||||
|
def test_spend_steps_200():
|
||||||
|
assert spend_steps(1000, 850, 200) == 0
|
||||||
|
assert spend_steps(1000, 800, 200) == 1
|
||||||
|
assert spend_steps(1000, 600, 200) == 2
|
||||||
|
assert spend_steps(1000, 1100, 200) == 0 # top-up direction
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_notifies_each_step():
|
||||||
|
state = BalanceWatchState(baseline_rub=1000.0)
|
||||||
|
state, msgs = plan_balance_notices(state, 750.0, step_rub=200.0)
|
||||||
|
assert state.last_notified_step == 1
|
||||||
|
assert len(msgs) == 1
|
||||||
|
assert "200" in msgs[0]
|
||||||
|
state, msgs2 = plan_balance_notices(state, 550.0, step_rub=200.0)
|
||||||
|
assert state.last_notified_step == 2
|
||||||
|
assert len(msgs2) == 1
|
||||||
|
# no re-notify same step
|
||||||
|
state, msgs3 = plan_balance_notices(state, 520.0, step_rub=200.0)
|
||||||
|
assert msgs3 == []
|
||||||
|
assert state.last_notified_step == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_topup_resets():
|
||||||
|
state = BalanceWatchState(baseline_rub=500.0, last_notified_step=2)
|
||||||
|
state, msgs = plan_balance_notices(state, 900.0, step_rub=200.0)
|
||||||
|
assert state.baseline_rub == 900.0
|
||||||
|
assert state.last_notified_step == 0
|
||||||
|
assert any("пополнен" in m for m in msgs)
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_low_once():
|
||||||
|
state = BalanceWatchState(baseline_rub=1000.0)
|
||||||
|
state, msgs = plan_balance_notices(state, 150.0, step_rub=200.0, low_rub=200.0)
|
||||||
|
assert state.low_notified
|
||||||
|
assert any("низкий" in m for m in msgs)
|
||||||
|
state, msgs2 = plan_balance_notices(state, 100.0, step_rub=200.0, low_rub=200.0)
|
||||||
|
assert not any("низкий" in m for m in msgs2)
|
||||||
@@ -0,0 +1,55 @@
|
|||||||
|
from gpu_rent.ready import ServiceCheck, _expected_services, verify_stack_local
|
||||||
|
|
||||||
|
|
||||||
|
class _Cfg:
|
||||||
|
enable_swarmui = True
|
||||||
|
llm_runtime = "none"
|
||||||
|
swarmui_local_port = 17801
|
||||||
|
ollama_local_port = 17811
|
||||||
|
llamacpp_local_port = 17812
|
||||||
|
|
||||||
|
|
||||||
|
def test_expected_services_swarm_only():
|
||||||
|
assert _expected_services(_Cfg()) == (True, False, False)
|
||||||
|
|
||||||
|
|
||||||
|
def test_expected_services_llm_only():
|
||||||
|
class C:
|
||||||
|
enable_swarmui = False
|
||||||
|
llm_runtime = "llamacpp"
|
||||||
|
|
||||||
|
assert _expected_services(C()) == (False, False, True)
|
||||||
|
|
||||||
|
|
||||||
|
def test_verify_stack_local_empty_when_nothing():
|
||||||
|
class C:
|
||||||
|
enable_swarmui = False
|
||||||
|
llm_runtime = "none"
|
||||||
|
swarmui_local_port = 17801
|
||||||
|
ollama_local_port = 17811
|
||||||
|
llamacpp_local_port = 17812
|
||||||
|
|
||||||
|
logs: list[str] = []
|
||||||
|
assert verify_stack_local(C(), logs.append, timeout=0.1) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_verify_stack_local_fails_closed_port(monkeypatch):
|
||||||
|
class C:
|
||||||
|
enable_swarmui = False
|
||||||
|
llm_runtime = "ollama"
|
||||||
|
swarmui_local_port = 17801
|
||||||
|
ollama_local_port = 17999
|
||||||
|
llamacpp_local_port = 17812
|
||||||
|
|
||||||
|
monkeypatch.setattr("gpu_rent.ready._tcp_ok", lambda port, host="127.0.0.1", timeout=0.8: False)
|
||||||
|
logs: list[str] = []
|
||||||
|
try:
|
||||||
|
verify_stack_local(C(), logs.append, timeout=0.3, poll_every=0.1)
|
||||||
|
assert False, "expected CloudError"
|
||||||
|
except Exception as exc:
|
||||||
|
assert "ollama" in str(exc).lower() or "не отвечает" in str(exc)
|
||||||
|
|
||||||
|
|
||||||
|
def test_service_check_dataclass():
|
||||||
|
c = ServiceCheck("x", True, "ok", "vm")
|
||||||
|
assert c.ok and c.where == "vm"
|
||||||
Reference in New Issue
Block a user