Your mean-reversion strategy was up 12% YTD. At 9:47 PM, a rare market microstructure event triggered a burst of malformed data. The Python process crashed. Nobody was watching. By the time you checked your phone the next morning, the strategy had been offline for 11 hours — and a short squeeze had already unwound the entire edge.
This is not a hypothetical. It is the most common failure mode in systematic trading: a process dies, nobody notices, and opportunity evaporates or losses accumulate. The fix is not a more robust strategy. It is a production-grade process supervision layer that detects failures instantly, restarts cleanly, and survives server reboots without manual intervention.
This article covers two production-vetted tools for exactly that: systemd (for system-level services) and supervisord (for user-space process management). We will walk through real configuration files, health check scripts, and the decision framework for choosing the right tool for each type of trading component.
1. The Problem: What Actually Kills Trading Processes
Before writing a single line of configuration, it helps to understand the specific failure modes that affect live trading systems.
Crash categories that matter:
| Failure type | Frequency | Recovery urgency |
|---|---|---|
| Unhandled exception in strategy logic | Medium | Immediate — positions may be left open |
| Data feed disconnect (WebSocket timeout) | High | Immediate — strategy operates blind |
| Memory leak leading to OOM kill | Low-Medium | Gradual — depends on leak rate |
| Out-of-disk on log volume | Low | Moderate — degraded logging masks other issues |
| Server reboot (patch, power event) | Low | Critical — full downtime window |
| Dependency service failure (Redis, DB) | Medium | Immediate — strategy cannot persist state |
Each failure type demands a slightly different response. The goal of a process supervisor is not just "restart when it dies." It is: restart fast, restart clean, alert someone, and preserve logs for post-mortem. That requires configuration discipline, not just defaults.
2. Choosing Your Tool: systemd vs. supervisor
Both tools solve the restart problem. They are not interchangeable.
| Dimension | systemd | supervisord |
|---|---|---|
| Scope | System-level, OS boot integration | User-space, no root required |
| Process model | One service = one unit file | One group = multiple worker processes |
| Restart policy granularity | Rich: on-failure, on-abnormal, backoff | Simpler: fixed restart delay, max retries |
| Dependency management | Built-in: After=, Requires=, Wants= |
None — ordering is manual |
| Log integration | Journalctl, structured logging | stdout/stderr to files, log rotation via external tool |
| Web UI | None natively | Built-in HTTP interface |
| Best for | Core services: data feed adapters, gateway processes | Strategy runners, one-off scripts, non-root deployments |
| Learning curve | Steeper — unit files, targets, environment masks | Gentle — INI-style config |
The practical rule: Use systemd for services that need to start at boot, run as root or a specific system user, or have hard dependencies on other services. Use supervisor for strategy processes, research-to-production handoff scripts, and environments where you do not have root access (e.g., shared research servers).
A typical production setup uses both: systemd manages the data feed adapter and database connection pool; supervisor manages individual strategy processes and the WebSocket heartbeat monitor.
3. systemd: System-Level Process Guardian
3.1 Anatomy of a Trading Service Unit File
A systemd unit file lives in /etc/systemd/system/ (system-wide) or ~/.config/systemd/user/ (per-user, for non-root deployments). For trading infrastructure, system-wide is almost always correct — you want the strategy to restart even if no one is logged in.
Here is a complete unit file for a live strategy process:
[Unit]
Description=Mean Reversion Strategy — US Equities
After=network-online.target
Wants=network-online.target
StartLimitIntervalSec=300
StartLimitBurst=5
[Service]
Type=simple
User=trading
Group=trading
WorkingDirectory=/opt/strategies/mean-reversion
Environment="PYTHONPATH=/opt/strategies/lib"
Environment="TICKDB_API_KEY=%{ENV:TICKDB_API_KEY}"
ExecStart=/opt/strategies/venv/bin/python -u main.py --mode live
Restart=on-failure
RestartSec=5
TimeoutStartSec=30
TimeoutStopSec=60
# Resource limits — critical for production
MemoryMax=2G
MemoryHigh=1.5G
LimitNOFILE=65536
# Logging
StandardOutput=journal
StandardError=journal
SyslogIdentifier=strategy-mean-reversion
# Security hardening
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/opt/strategies/logs
PrivateTmp=true
[Install]
WantedBy=multi-user.target
Key decisions explained:
Type=simple: The process does not fork. systemd monitors the main PID directly. For Python strategies that stay in the foreground, this is correct.Restart=on-failure: Restarts only on abnormal exit (non-zero exit code, signal-induced death, watchdog timeout). Does not restart on cleansys.exit(0).StartLimitBurst=5withinStartLimitIntervalSec=300: After 5 restarts within 5 minutes, systemd stops attempting to restart and marks the service as failed. This prevents a crash loop from consuming all CPU.MemoryMax=2G: Hard cap. Prevents a memory leak from consuming the entire system. This alone has saved production systems multiple times.ProtectSystem=strict: Mounts/usr,/boot,/etcread-only. Prevents a compromised strategy process from modifying system binaries.
3.2 Reload and Manage
# Reload systemd after editing a unit file
sudo systemctl daemon-reload
# Start, stop, restart
sudo systemctl start strategy-mean-reversion
sudo systemctl restart strategy-mean-reversion
sudo systemctl stop strategy-mean-reversion
# Check status
sudo systemctl status strategy-mean-reversion
# View logs (last 100 lines, follow mode)
sudo journalctl -u strategy-mean-reversion --no-pager -n 100 -f
# View only restarts
sudo journalctl -u strategy-mean-reversion --no-pager -g "service: main" --since "1 hour ago"
3.3 Dependency Chaining: Ensuring Your Data Feed Starts First
A strategy that starts before its data feed is worse than a crashed strategy — it will trade on stale data or fail with a confusing connection error. Use systemd's dependency directives to enforce ordering:
[Unit]
Description=Market Data Feed Adapter
After=network-online.target
Wants=network-online.target
[Service]
ExecStart=/opt/strategies/bin/data-feed-adapter
Restart=always
RestartSec=3
# ... other fields
Then, in your strategy unit file:
[Unit]
Description=Mean Reversion Strategy — US Equities
After=network-online.target market-data-feed.service
Requires=market-data-feed.service
The Requires= directive means: if market-data-feed.service stops, this strategy also stops. This is the correct behavior — you do not want your strategy running against an unavailable data feed.
4. supervisor: User-Space Process Manager
supervisord is the right choice when you need a lightweight process manager without root access, want a built-in web UI for non-technical team members, or need to manage multiple strategy processes under a single supervisor instance.
4.1 Installing and Starting supervisor
# Ubuntu / Debian
sudo apt-get install supervisor
# CentOS / RHEL
sudo yum install supervisor
# Start and enable on boot
sudo systemctl enable supervisord
sudo systemctl start supervisord
The main configuration file lives at /etc/supervisor/supervisord.conf. Per-program configuration goes into /etc/supervisor/conf.d/.
4.2 Strategy Process Configuration
[program:strategy-mean-reversion]
command=/opt/strategies/venv/bin/python -u /opt/strategies/mean-reversion/main.py --mode live
directory=/opt/strategies/mean-reversion
user=trading
autostart=true
autorestart=true
startretries=5
exitcodes=0;2;SIGHUP;SIGTERM
startsecs=10
stopwaitsecs=60
stdout_logfile=/opt/strategies/logs/strategy-mean-reversion-stdout.log
stdout_logfile_maxbytes=100MB
stdout_logfile_backups=5
stderr_logfile=/opt/strategies/logs/strategy-mean-reversion-stderr.log
stderr_logfile_maxbytes=100MB
stderr_logfile_backups=5
environment=PYTHONPATH="/opt/strategies/lib",TICKDB_API_KEY="%(ENV_TICKDB_API_KEY)s"
priority=998
Key decisions explained:
autorestart=true: Restart on any exit, including unexpected crashes. Combined withexitcodes=0;2;SIGHUP;SIGTERM, supervisord distinguishes between "expected" exits (clean shutdown) and failures.startretries=5: After 5 consecutive failures, supervisor stops attempting and requires a manualsupervisorctl reread && supervisorctl update.startsecs=10: supervisor waits 10 seconds before marking the process as successfully started. If the process crashes within 10 seconds, it counts as a failure. This prevents false positives where a process spawns but immediately dies.%(ENV_TICKDB_API_KEY)s: supervisor supports environment variable interpolation. Your API key lives in the system environment and is injected at runtime — never hardcoded in a config file.priority=998: Lower numbers start first. If you have a data feed adapter atpriority=900and this strategy atpriority=998, the feed starts before the strategy.
4.3 The Web Interface
supervisord ships with an HTTP server. Add this to your supervisord.conf:
[inet_http_server]
port=*:9001
username=admin
password=<hashed-password>
[supervisorctl]
serverurl=http://localhost:9001
Generate a password hash:
echo -n "yourpassword" | sha256sum
Then use supervisorctl to manage processes without SSH:
# Status of all programs
supervisorctl status
# Restart a specific strategy
supervisorctl restart strategy-mean-reversion
# Tail logs
supervisorctl tail -f strategy-mean-reversion stderr
# Add a new program without restarting supervisord
supervisorctl reread
supervisorctl add strategy-new-strategy
4.4 Process Groups: Managing Multiple Strategies
For a multi-strategy deployment, organize processes into groups:
[group:strategies]
programs=strategy-mean-reversion,strategy-momentum,strategy-arbitrage
[program:strategy-mean-reversion]
; ... config ...
[program:strategy-momentum]
command=/opt/strategies/venv/bin/python -u /opt/strategies/momentum/main.py --mode live
directory=/opt/strategies/momentum
; ...
With groups, you can control all strategies at once:
supervisorctl restart strategies/*
supervisorctl stop strategies/*
supervisorctl status strategies/*
5. Health Check Scripts: The Supervisor Between the Supervisor and Your Process
Both systemd and supervisor can restart a process, but neither intrinsically knows whether the process is healthy — only whether it is running. A crashed process that loops into a non-functional state (e.g., connected to a dead WebSocket feed) will pass a restart check but produce garbage signals.
This is where a health check script bridges the gap.
5.1 The Strategy Health Check Script
#!/usr/bin/env bash
# health_check.sh — Place in /opt/strategies/bin/health_check.sh
# Usage: health_check.sh <strategy_name> <pid_file> <max_missing_heartbeats>
set -euo pipefail
STRATEGY_NAME="${1:-}"
PID_FILE="${2:-/tmp/strategy.pid}"
MAX_MISSING_HEARTBEATS="${3:-3}"
HEARTBEAT_FILE="/tmp/heartbeat_${STRATEGY_NAME}.tsv"
# Verify the process is actually running
if [[ ! -f "$PID_FILE" ]]; then
echo "[$(date -u)] ALERT: No PID file for $STRATEGY_NAME"
exit 1
fi
PID=$(cat "$PID_FILE")
if ! kill -0 "$PID" 2>/dev/null; then
echo "[$(date -u)] ALERT: Process $PID is not running"
exit 1
fi
# Check heartbeat file
if [[ -f "$HEARTBEAT_FILE" ]]; then
LAST_HEARTBEAT=$(stat -c %Y "$HEARTBEAT_FILE" 2>/dev/null || echo "0")
NOW=$(date +%s)
MISSING_SECONDS=$((NOW - LAST_HEARTBEAT))
# Heartbeat should arrive every 5 seconds from the strategy
MAX_INTERVAL=$((MAX_MISSING_HEARTBEATS * 5))
if [[ $MISSING_SECONDS -gt $MAX_INTERVAL ]]; then
echo "[$(date -u)] ALERT: No heartbeat for ${MISSING_SECONDS}s (max: ${MAX_INTERVAL}s)"
kill -9 "$PID" 2>/dev/null || true
exit 1
fi
else
echo "[$(date -u)] WARN: No heartbeat file yet for $STRATEGY_NAME"
fi
# Check for open positions flagged as stale (dead letter mechanism)
if [[ -f "/tmp/${STRATEGY_NAME}_stale_position.lock" ]]; then
echo "[$(date -u)] ALERT: Stale position lock detected — forcing restart"
kill -9 "$PID" 2>/dev/null || true
exit 1
fi
echo "[$(date -u)] OK: $STRATEGY_NAME (PID $PID) is healthy"
exit 0
5.2 Integrating Health Checks with systemd
systemd natively supports health checks via ExecStartPost and TimeoutStartSec, but for a more robust check, use a systemd service that wraps the health check:
[Unit]
Description=Health Check for Mean Reversion Strategy
After=strategy-mean-reversion.service
Requires=strategy-mean-reversion.service
[Timer]
OnBootSec=30
OnUnitActiveSec=15
Unit=health-check-mean-reversion.service
[Install]
WantedBy=timers.target
And the companion service unit:
[Unit]
Description=Health Check Runner
[Service]
Type=oneshot
ExecStart=/opt/strategies/bin/health_check.sh mean-reversion /tmp/strategy-mean-reversion.pid 3
The timer runs the health check every 15 seconds after the first boot delay. If the health check exits non-zero, it triggers an alert (you can extend the script to send a Slack message, PagerDuty alert, or kill the process and let systemd's Restart=on-failure handle the recovery).
5.3 Integrating Health Checks with supervisor
supervisor does not natively run health checks, but you can add a dedicated program that does:
[program:health-check-mean-reversion]
command=/opt/strategies/bin/health_check.sh mean-reversion /tmp/strategy-mean-reversion.pid 3
autostart=true
autorestart=true
startsecs=5
numprocs=1
process_name=%(program_name)s
stdout_logfile=/opt/strategies/logs/health-check-mean-reversion.log
stderr_logfile=/opt/strategies/logs/health-check-mean-reversion-stderr.log
The health check process itself stays alive. If it detects a stale strategy, it exits with code 1, supervisor interprets that as a failure, and supervisor restarts the health check. A separate execrestart-on-alert.sh script (triggered from the health check) kills and restarts the strategy process.
6. Inside the Strategy: Heartbeat Emission
A health check script is only as good as the heartbeat the strategy emits. Here is the production pattern for embedding a heartbeat into a Python strategy:
import os
import time
import signal
import sys
import threading
class HeartbeatManager:
def __init__(self, strategy_name: str, interval: int = 5):
self.strategy_name = strategy_name
self.interval = interval
self.heartbeat_path = f"/tmp/heartbeat_{strategy_name}.tsv"
self.running = False
self._thread = None
def _emit_heartbeat(self) -> None:
"""Touch the heartbeat file. Health check uses mtime as the timestamp."""
open(self.heartbeat_path, "w").close()
# Also write structured data for debugging
with open(self.heartbeat_path, "w") as f:
f.write(f"ts={time.time()}\tpid={os.getpid()}\tstrategy={self.strategy_name}\n")
def _loop(self) -> None:
while self.running:
self._emit_heartbeat()
time.sleep(self.interval)
def start(self) -> None:
self.running = True
self._thread = threading.Thread(target=self._loop, daemon=True)
self._thread.start()
def stop(self) -> None:
self.running = False
if self._thread:
self._thread.join(timeout=2)
try:
os.remove(self.heartbeat_path)
except FileNotFoundError:
pass
# Graceful shutdown handling
heartbeat = HeartbeatManager("mean-reversion", interval=5)
def shutdown(signum, frame):
print("Received shutdown signal — stopping heartbeat and exiting")
heartbeat.stop()
sys.exit(0)
signal.signal(signal.SIGTERM, shutdown)
signal.signal(signal.SIGINT, shutdown)
heartbeat.start()
# --- Your strategy logic runs here ---
try:
run_strategy()
finally:
heartbeat.stop()
The heartbeat file uses the file's modification time as the clock. The health check script calls stat -c %Y to read it — no parsing required, no IPC overhead. If the strategy hangs (e.g., in a deadlocked network call), the heartbeat stops, and the health check detects it within 3 × 5 = 15 seconds.
7. Server Reboot: The Critical Missing Piece
Process supervisors restart crashed processes. They do not inherently survive a server reboot — unless configured correctly.
For systemd:
The WantedBy=multi-user.target in the [Install] section links the service to a boot target. When the server starts, systemd activates multi-user.target, which starts all services that declare WantedBy=multi-user.target. This is automatic and requires no additional configuration.
Verify it works:
sudo systemctl is-enabled strategy-mean-reversion
# Expected output: generated
For supervisor:
supervisord itself must be started at boot. The installed package registers itself as a systemd service automatically on modern distributions:
sudo systemctl is-enabled supervisord
# Expected output: enabled
Once supervisord is running, it reads its configuration at startup and launches all programs marked autostart=true. There is no separate per-program enable step — every program with autostart=true starts on boot.
Critical: Ensure data feed starts before strategy on reboot.
# In the strategy program config
priority=998 # Higher number = starts later
And in systemd, use the dependency chain described in Section 3.3. A 30-second delay on the strategy startup after the data feed, combined with the health check, gives the data adapter time to establish its connections before the strategy begins evaluating signals.
8. Alerting: Knowing Something Went Wrong
Restarting a process is necessary but not sufficient. You need to know that it restarted, and why. Both tools integrate with standard alerting systems.
systemd + systemdd + a simple alerting script:
Extend your health check script to send alerts:
# Add to health_check.sh after the stale-position check
if [[ $? -ne 0 ]]; then
curl -s -X POST "https://hooks.example.com/alert" \
-H "Content-Type: application/json" \
-d "{\"text\":\"[ALERT] Strategy $STRATEGY_NAME unhealthy — restarting\"}"
fi
systemd status change alerts:
Use a systemd path unit to watch for status changes:
# Monitor via journalctl in a separate daemon process
journalctl -f -u strategy-mean-reversion -o json | jq -r 'select(.SYSTEMD_UNIT=="strategy-mean-reversion.service" and .SYSLOG_IDENTIFIER=="strategy-mean-reversion") | .MESSAGE' | while read msg; do
echo "[$(date)] $msg"
done
supervisor event listeners:
supervisord has a native event notification system. Define a listener program:
[eventlistener:alerts]
command=/opt/strategies/bin/supervisor_alert_listener.py
events=PROCESS_STATE_EXITED,PROCESS_STATE_FATAL
buffer_size=100
stdout_logfile=/opt/strategies/logs/supervisor-alerts.log
stderr_logfile=/opt/strategies/logs/supervisor-alerts-stderr.log
autorestart=true
The listener script receives JSON event payloads when any supervised process changes state. This is where you plug in Slack notifications, PagerDuty, or email alerts.
9. Decision Framework and Deployment Summary
Use the following decision tree when deploying a new trading component:
Is the component a core system service that must start at boot
and has hard dependencies on other services (data feed, database)?
├── YES → Use systemd
│ ├── Define dependency chain with After=, Requires=
│ ├── Set MemoryMax to prevent resource exhaustion
│ ├── Enable StartLimitBurst to prevent crash loops
│ └── Use a separate health-check timer unit
└── NO → Does the deployment run as a non-root user
or need a web UI for operational visibility?
├── YES → Use supervisor
│ ├── Group related processes with [group:]
│ ├── Configure priority for startup ordering
│ └── Add event listener for alerts
└── NO → Use systemd (preferred for reliability)
Both tools require discipline in configuration. The most common mistake is treating the default settings as production-ready. They are not. Memory limits, restart backoff, health checks, log rotation, and alerting all need explicit configuration.
10. Closing
A trading strategy without process supervision is a single point of failure with no recovery path. The market does not pause because your process died. The spread does not wait while you SSH in to restart.
systemd and supervisor are not competing choices — they are complementary layers. systemd owns the infrastructure layer (data feeds, gateways, network interfaces). supervisor owns the strategy layer (individual strategy processes, research handoffs, one-off execution scripts). A health check script bridges the two, detecting the failure modes that a simple process-running check misses: deadlocks, stale feeds, and orphaned positions.
The configuration files in this article are production-ready starting points. Copy them, adjust the paths and thresholds, and treat the health check script as the first thing you write for any new strategy — not an afterthought.
Next Steps
If you're building your first live trading infrastructure: Start with supervisor for your strategy processes. It is easier to debug and has a built-in web UI for operational monitoring. Add systemd for your data feed adapter.
If you're migrating from research to production: Place your health check script in /opt/strategies/bin/ alongside your strategy code. Treat it as part of the deployment package, not a separate operational concern.
If you need institutional-grade historical data for backtesting before your strategy goes live: Sign up at tickdb.ai for access to 10+ years of cleaned US equity OHLCV data via the /v1/market/kline endpoint — no credit card required for the free tier.
If you're deploying across multiple servers and need a unified process monitoring dashboard: Look into supervisord with the supervisorweb plugin or evaluate systemd's systemctl status aggregation via systemd-run and remote journal forwarding.
This article does not constitute investment advice. Markets involve risk; past performance does not guarantee future results. Always conduct thorough out-of-sample testing before deploying any strategy in a live environment.