Production deployment checklist
A single page to walk through before pointing proxxx at a real cluster you care about. Each item is a one-line check + a verifiable command. Treat this as the minimum bar — your shop's own runbook may be stricter.
TIP
This page is for the operator deploying proxxx, not for the PVE cluster itself. Cluster hardening (corosync over a private ring, firewall rules, certificate rotation) is upstream Proxmox material; we link to it but don't restate it.
On-call essential:
incident freezeis a fleet-wide write kill-switch — during a maintenance window or a suspected compromise, every mutation is refused before it reaches Proxmox until youthaw.
1. Verify the binary
[ ] Download from a tagged release, not main
TARGET=x86_64-unknown-linux-musl # or aarch64-apple-darwin
VERSION=$(gh release view --repo fabriziosalmi/proxxx --json tagName -q .tagName | sed 's/^v//') # latest tag
gh release download v${VERSION} \
--repo fabriziosalmi/proxxx \
--pattern "*-${TARGET}.tar.gz" \
--pattern "*-${TARGET}.tar.gz.sha256"[ ] Check the SHA-256 sidecar
shasum -a 256 -c proxxx-${VERSION}-${TARGET}.tar.gz.sha256
# → proxxx-...tar.gz: OK[ ] Verify the sigstore keyless cosign signature (release ≥ next-tag-after-v0.1.6)
gh release download v${VERSION} \
--repo fabriziosalmi/proxxx \
--pattern "*-${TARGET}.tar.gz.cosign.bundle"
cosign verify-blob \
--bundle proxxx-${VERSION}-${TARGET}.tar.gz.cosign.bundle \
--certificate-identity-regexp 'https://github.com/fabriziosalmi/proxxx/.github/workflows/release.yml@.*' \
--certificate-oidc-issuer 'https://token.actions.githubusercontent.com' \
proxxx-${VERSION}-${TARGET}.tar.gz
# → Verified OKThe cert-identity-regexp pins the OIDC subject to this exact workflow path in this exact repo — a leaked sigstore cert from any other workflow or any other repo can't validate against these bundles. The transparency-log inclusion proof is embedded in the bundle, so verification is offline.
[ ] Audit the CycloneDX SBOM (optional but recommended)
gh release download v${VERSION} --repo fabriziosalmi/proxxx \
--pattern "*.cdx.json" --pattern "*.cdx.json.sha256"
shasum -a 256 -c proxxx-${VERSION}.cdx.json.sha256
grype sbom:proxxx-${VERSION}.cdx.json # or trivy / cyclonedx-cli2. Configure access
[ ] Use API tokens, not passwords
Tokens are revocable, scopable, and don't carry full account privilege when --privsep=1. Create with:
ssh root@<node>
pveum user token add operator@pve proxxx --privsep=1
# Grants the TOKEN the same role you grant on the user side:
pveum acl modify /vms/100 -tokens 'operator@pve!proxxx' -roles PVEVMAdmin[ ] Pin verify_tls = true unless you know why not
Self-signed labs flip this to false for convenience. In production:
verify_tls = trueIf you're running PVE behind a real cert (Let's Encrypt, internal CA, ACME via PVE itself), this is the only correct setting. Disabling TLS verification exposes the entire API + WebSocket traffic (including serial-console tickets) to any MITM on the path.
[ ] Store the token secret in the OS keychain or a 0600 file, not inline
# Option A: macOS keychain
security add-generic-password \
-a "$USER" -s proxxx -w "<token-uuid>"
# Option B: Linux secret-service (gnome-keyring / kwallet)
secret-tool store --label proxxx service proxxx account token_secret
# (proxxx reads via the `keyring` crate, falling back to
# secret-service on Linux, libsecret on macOS keychain)
# Option C: 0600 file referenced from config.toml
mkdir -p ~/.config/proxxx
printf '%s' '<uuid>' > ~/.config/proxxx/token.secret
chmod 600 ~/.config/proxxx/token.secret
# token_secret_file = "~/.config/proxxx/token.secret" in config.tomlWARNING
proxxx refuses to read token_secret_file if the file is not mode 0600 (Unix). It will print Security Error: token_secret_file '<path>' has unsafe permissions <mode> and exit. This is intentional — don't chmod 644 to "fix" it.
[ ] Validate the connection works as the deploying user
proxxx ls nodes
proxxx ls guests --format json | jq '.[] | {vmid, name, status}'
proxxx perms <user>3. Configure HITL (if any operator runs destructive ops)
[ ] Provision a dedicated Telegram bot
Don't reuse a bot that's also wired to other systems — the HITL daemon polls and acknowledges every callback, and a shared bot's other listeners may double-fire.
# In Telegram: chat with @BotFather
/newbot → name + username → copy the API token
/setprivacy → DISABLE (so the bot sees group messages)[ ] Pin [telegram] in config.toml
[telegram]
bot_token = "<bot-api-token>"
chat_id = "<your-numeric-chat-id>"The bot token resolves with the same hierarchy as the PVE token: PROXXX_TELEGRAM_BOT_TOKEN env, bot_token_file, keychain, inline.
[ ] Set allowed_approvers — without it the daemon refuses every callback
The callback signature proves proxxx minted the approval keyboard. It does not establish who pressed the button: Telegram delivers the buttons to everyone who can see the message. Without an allowlist, any member of chat_id — including someone added to the group later — could approve a destructive operation, which then runs with the daemon's full PVE credentials.
[telegram]
chat_id = "-1001234567890"
allowed_approvers = [123456789, 987654321] # numeric ids, not usernamesGet each approver's numeric id by having them message @userinfobot. Usernames are deliberately not accepted: a Telegram handle can be released and re-registered by someone else, so an allowlist keyed on handles is forgeable.
An absent or empty list is treated as "refuse everything" rather than "allow anyone" — the same fail-closed posture as a destructive MCP tool with no matching policy.
[ ] Configure [[policies]] rules
[[policies]]
when = { action = "delete", tag = "prod" }
require = "telegram-2of3" # or "telegram"
channel = "telegram"
[[policies]]
when = { action = "stop", vmid = "100" }
require = "telegram"
channel = "telegram"Policies match by action, tag, vmid, or wildcard. The deny-on-timeout is hardcoded to 120 s — if the human doesn't approve in that window, the op is rejected (NOT auto-approved).
[ ] Run the HITL daemon under a process supervisor
# systemd unit at /etc/systemd/system/proxxx-hitl.service:
[Unit]
Description=proxxx HITL approval daemon
After=network-online.target
[Service]
Type=simple
User=proxxx-ops
ExecStart=/usr/local/bin/proxxx hitl serve
Restart=on-failure
RestartSec=5
# Log records go to stderr as well as the rotating file, so
# `journalctl -u proxxx-hitl` shows warnings and errors — including
# the freeze lock becoming unreadable and TLS pinning being skipped.
# Raise verbosity per-crate when debugging:
# Environment=RUST_LOG=proxxx=debug,russh=debug
Environment=RUST_LOG=proxxx=info
# NOTE: replay protection is session-local — a restart clears
# the consumed-txn-id set, so an approval callback that was
# already used becomes usable again until the keyboard is
# superseded. Approver authorisation (allowed_approvers) and
# the per-request txn_id nonce are the controls that survive
# a restart. See src/hitl/pending.rs ("Scope honesty").
[Install]
WantedBy=multi-user.target[ ] Test the round-trip end-to-end
proxxx hitl test --action delete --vmid 999
# → Telegram → tap Approve in <120s → daemon runs the op
# → daemon answers callback "✅ Done" → message lifecycle done4. Configure alerting (optional)
[ ] Define [[alerts]] rules in config.toml
[[alerts]]
name = "node_offline"
when = "node_offline"
for_secs = 120
severity = "critical"
route = ["telegram", "ntfy:proxxx-prod"]
dedup_secs = 600Predicates available: node_offline, storage_above, replication_failing. The dedup_secs window prevents re-fire spam.
[ ] Run the alert daemon under a supervisor
ExecStart=/usr/local/bin/proxxx alerts watch --interval 30The daemon persists its dedup window to SQLite (cache schema 1 → 2 since v0.1.2), so a routine restart doesn't re-fire every active alert. Persistence is local to the daemon's host — no shared state across replicas.
[ ] Test each route once
proxxx alerts test --route 'telegram'
proxxx alerts test --route 'ntfy:proxxx-prod'
proxxx alerts test --route 'webhook:https://hooks.example/notify'5. Configure SSH layer (if running proxxx perms or proxxx patch apply)
[ ] Provision a dedicated SSH key for proxxx
Don't reuse your personal ~/.ssh/id_ed25519. proxxx maintains its own known_hosts at $XDG_CONFIG_HOME/proxxx/known_hosts — giving it a dedicated key keeps the audit trail separate.
ssh-keygen -t ed25519 -f ~/.ssh/proxxx_ops -N "" \
-C "proxxx-ops@$(hostname)"
ssh-copy-id -i ~/.ssh/proxxx_ops.pub root@<each-node>[ ] Configure [ssh] block
[ssh]
user = "root"
key_path = "~/.ssh/proxxx_ops"
strict_host_key_checking = "tofu" # default; pinning happens on first connect[ ] Verify the round-trip
proxxx perms <user> # exercises the same path
# → table of effective ACLs; if you see the table, SSH works.6. Lock down the operator host itself
[ ] Treat ~/.config/proxxx/config.toml as a credential file
chmod 600 ~/.config/proxxx/config.toml
ls -l ~/.config/proxxx/[ ] Confirm shells aren't leaking secrets via history
# Bash:
echo $HISTFILE
grep -E 'PROXXX_TOKEN_SECRET|PROXXX_PBS_TOKEN_SECRET' ~/.bash_history
# Zsh: ~/.zsh_history. Fish: ~/.local/share/fish/fish_history.If you find tokens in history, rotate them on the cluster side before deleting from history.
[ ] Pin a Rust version + build from source for reproducibility (optional)
git clone https://github.com/fabriziosalmi/proxxx.git
cd proxxx
git checkout v${VERSION}
cargo build --release --target x86_64-unknown-linux-musl
# Compare your binary's sha256 against the release sha256.7. Operational runbook
[ ] Pin the build into your fleet inventory
proxxx version --json
# → { "version": "0.1.6", "git_sha_short": "...", "audit_ignores_count": 1, ... }Snapshot this output into your inventory (Ansible facts / Salt grain / Puppet fact) so a security advisory triage can answer "which hosts have <vulnerable version>" instantly.
[ ] Subscribe to release notifications
GitHub repo → Watch → Custom → Releases. Or pin the release feed: https://github.com/fabriziosalmi/proxxx/releases.atom.
[ ] Document your local risk-override policy
--allow-risk bypasses the pre-flight gate. If your shop permits this for any class of op (e.g. patch-apply during a planned window), document who can use it and for what in your runbook. The flag is ungated by design — proxxx trusts the operator who typed --yes AND --allow-risk.
[ ] Test recovery: token revocation
ssh root@<node>
pveum user token remove operator@pve proxxx
# → next proxxx call from the operator host should 401
proxxx ls nodes # expect: HTTP 401 No ticketIf it doesn't 401, the token wasn't actually scoped — re-issue with --privsep=1 and grant the role to the token path.
8. Harden the MCP server (if you expose proxxx over MCP)
[ ] Keep the transport on loopback, or set mcp_token before exposing it
proxxx mcp serve-http refuses to start on a non-loopback bind (e.g. 0.0.0.0) unless mcp_token is set:
# Local-only (default, safe): binds 127.0.0.1
proxxx mcp serve-http
# Network-exposed without a token → refused
proxxx mcp serve-http --bind 0.0.0.0:8080
# → error: refusing to bind a non-loopback address without a token.
# Set `mcp_token` (or pass --token), bind to 127.0.0.1, or pass
# --insecure-bind to override.mcp_token is a profile config field (or the --token flag). An empty or whitespace token counts as absent — don't set it to "" and think you're covered. When the server is network-exposed and the token is absent, every request is denied (fail-closed), and that denial survives a SIGHUP that clears the token from a live config.
--insecure-bind overrides the refusal. Never pass it on an untrusted network — it exists for consciously-chosen trusted-segment cases only (see AR-3 in ACCEPTED-RISKS.md).
[ ] Gate every destructive MCP tool with a [[policies]] entry
Over MCP, proxxx is fail-closed: a destructive tool with no matching [[policies]] entry is refused (a typed isError envelope — the PVE gateway is never reached). There is no ungated inline path. The only way to run a destructive tool over MCP is a matching policy that routes it to HITL approval.
Destructive (policy-gated) tools: stop_guest, restart_guest, suspend_guest, delete_guest, create_guest, clone_guest, clone_with_cloudinit, migrate_guest, create_snapshot, delete_snapshot. Non-destructive tools stay inline: start_guest, resume_guest, and every get_* / list_* read.
# Without a matching policy, a destructive tool is simply unavailable
# to the MCP client. Confirm your rules cover every destructive action
# you intend a client to perform.
[[policies]]
when = { action = "delete" }
require = "telegram"
channel = "telegram"9. Guard unmanned auto-converge (if you run reconcile watch --auto-converge)
[ ] Set a small max_unmanned_changes before enabling converge_prune
converge_prune = true lets unmanned converge delete drifted resources — but it holds all deletes unless max_unmanned_changes (a per-tick change cap) is also set. With prune on and no cap, deletes are held and only create/update operations converge.
[reconcile]
auto_converge = true
converge_prune = true
max_unmanned_changes = 3 # keep this single-digit
allowed_families = ["lxc"] # whitelist which families converge unmannedSize the cap small — single digits. A large cap (e.g. 49) re-opens the 10–49 unmanned-delete band, which is exactly the blast radius the cap exists to close (see AR-4). Use allowed_families to whitelist which resource families are allowed to converge without a human.
10. Protect the audit chain
[ ] Keep audit.key at 0600, and consider a separate volume
The audit log is HMAC-signed. proxxx creates the audit dir 0700 and refuses to load a group- or world-readable audit.key (Unix), so keep it owner-only:
chmod 600 <audit-dir>/audit.key
ls -l <audit-dir>/audit.key # → -rw------- (0600)To keep the signing key off the same volume as the database it signs, point PROXXX_AUDIT_KEY at a separate volume — a same-user/root attacker who can rewrite the DB then can't silently re-sign it (AR-5). PROXXX_AUDIT_DIR relocates the DB and key together if you'd rather move the whole directory.
[ ] Run audit verify from a separate trust context
proxxx audit verify # exits non-zero if the chain has been tampered withRun this from a host or account that can't write the audit DB — a verifier that shares the operator's write access can't prove much. Wire the non-zero exit into your monitoring.
[ ] Rotate the audit key after a compromise — it no longer costs you the history
The HMAC key is 32 bytes beside the database it protects, so anyone who can read that file can forge the chain. If the host is ever compromised, rotate:
proxxx audit rotate-key # archives the old key, starts signing with a new one
proxxx audit verify # every entry still verifies, old and newUntil v0.13.4 rotating meant deleting the key and starting a new chain, which destroyed the verifiability of exactly the history an investigation needs — so in practice the key was permanent. Each row now records which key signed it, and retired keys are archived beside the primary as audit.key.<id>.
Back up the retired keys. They are part of the trail: without one, the rows it signed can no longer be verified, and proxxx audit verify counts them as failures rather than skipping them.
11. Upgrade, rollback and backup
[ ] Know what must survive the host
proxxx keeps two kinds of state in the platform data directory, and only one of them matters if the machine is rebuilt.
| Path | Must be preserved? | Why |
|---|---|---|
audit.db | Yes | The only record of who issued which mutation. Not reconstructible from anything else. |
audit.key | Yes | 32 bytes. Without it no surviving copy of audit.db can be verified — losing the key alone is enough to make the trail worthless. |
audit.key.retired-* | Yes | Keys retired by audit rotate-key. Each still proves the rows it signed; losing one turns those rows into verification failures. |
freeze.lock, freeze.<profile>.lock | No | Runtime kill-switch state. Absent means thawed, which is the correct default after a rebuild. |
cache.db | No | Cluster snapshots and the operation queue. Regenerates from the cluster on next start. |
proxxx.log* | No | 14 daily rotations, forensic convenience only. |
Config lives separately under the config directory and is recreatable with proxxx init, though backing it up saves re-entering the profile.
Point PROXXX_AUDIT_DIR at a volume that is already backed up if you would rather not add a new backup target:
# systemd unit
Environment=PROXXX_AUDIT_DIR=/var/lib/proxxx/audit[ ] Upgrade
systemctl stop proxxx-hitl # let in-flight approvals settle
# verify + install the new binary (section 1)
systemctl start proxxx-hitl
proxxx doctor # confirms config, auth and audit chainStopping first is deliberate: an approval that is parked when the process is replaced is lost from the daemon's in-memory replay window, and a request approved during the swap has nothing listening for the callback.
What is compatible across an upgrade:
- Config — backwards compatible. New keys default; old keys keep working.
- Audit DB — backwards compatible. v1 rows keep verifying under the v1 formula while new rows are written as v2.
- Declared state TOML — carries
meta.schema_version. A document newer than the binary understands is refused rather than reinterpreted. - Cache DB — not backwards compatible; see rollback.
[ ] Rollback
systemctl stop proxxx-hitl
# reinstall the previous binary
rm -f "$(proxxx doctor --json | jq -r '.cache_path // empty')" # or delete cache.db by hand
systemctl start proxxx-hitlThe cache database records a schema version and refuses to open when that version is newer than the running binary — an older proxxx will report cache DB schema version N is newer than this binary's M rather than silently misreading it. Deleting cache.db is safe: it is regenerated from the cluster on the next start. The audit DB and its key need no action, and must not be deleted.
See also
- Configuration schema — every TOML block by section.
- Security model — threat model + invariants.
- Pre-commit gate — what every release passes before tagging.
ACCEPTED-RISKS.md— residual risks (AR-1…AR-6) knowingly accepted for this release, with the guardrails above cross-referenced.- Troubleshooting — error message → fix index.
SECURITY.md— coordinated disclosure contact.