Engineering Notes · agent operations · observability

Logs Without the SSH Session

Lucas van Staden · ProxiBlue · October 2026 · proxiblue.com.au

My AI agents are not allowed to SSH to a live shop. Fine, except the answer to half the questions I ask them is sitting in an nginx or PHP-FPM log on that box. So the logs come to the agent instead. Hypernode's Supervisor feature runs a small shipper on the node, a self-hosted VictoriaLogs instance holds the data, and an MCP server gives the agent read-only, per-shop search.

Why bother

I already ship exceptions to Bugsink, and that covers the stack traces. It does not cover the rest of what you look at when a shop misbehaves: the burst of 502s at 3am, the one category page taking 12 seconds, the slow query log filling up after a reindex, the bot hammering the search endpoint. That all lives in plain log files on the Hypernode.

Reading those files used to mean someone SSHing in. I don't want agents doing that on live (read-only is a promise, not a guarantee, and one stray command is all it takes), and I don't want to be the human copy-paste layer between the server and the agent either. So the goal here is simple: get the logs off the box, keep them per shop and per environment, and let the agent query them through a tool that physically cannot write anything.

The shape of it

HYPERNODE NODE · PUSH-ONLY, NO ROOT LOG HOST · SMALL VPS DEV MACHINE log files nginx access · error php-fpm · php-slow mysql-slow magento var/log/*.log Vector parse · redact · disk buffer kept alive by Supervisor tails VictoriaLogs one tenant per shop + env 30 day retention hard disk cap Caddy · the only gate token → exactly one tenant tenant headers overwritten anything else → 403 tenant set write token · /insert/* only AI agent per-project MCP config brings its own read token uat token sees uat only mcp-victorialogs read-only query tools holds no tokens, no data passes the token through LogsQL read token · /select/logsql only no ssh to live
Both paths go through the same gate. The shop pushes with a token that can only write, the agent asks with a token that can only read, and Caddy decides the tenant for both. The dashed line is the shortcut the agent never gets.

Everything on the shop side is push-only. Nothing connects in to the Hypernode, and nothing on the node needs root.

The Hypernode bit: Supervisor

The piece that makes this possible is a Hypernode feature a lot of people never turn on. Hypernode ships Supervisor for long running programs, it watches whatever you give it and restarts it if it dies. It is off by default, and you turn it on with the hypernode-systemctl tool:

hypernode-systemctl settings supervisor_enabled
hypernode-systemctl settings supervisor_enabled True

(The first line just tells you the current value, worth running before you change anything.) If you deploy with Hypernode Deploy, you can keep this in source control instead, it has a HypernodeSettingConfiguration and a SupervisorConfiguration that set the flag and install your program configs on every deploy.

Programs are declared as *.conf files in /data/web/supervisor/, then loaded with supervisorctl:

supervisorctl reread
supervisorctl add vector
supervisorctl status

Supervisor's own logs end up in /data/var/log/supervisor/, which is the first place to look if something will not stay up.

What is on the box to ship

Hypernode is quite good here, most of the useful logs are already readable by the app user and some are already structured. These are the ones I ship (check them on your own node first, plans and setups differ):

LogPathNotes
nginx access/var/log/nginx/access.logAlready JSON, one object per line (Hypernode docs). No grok patterns needed.
nginx error/var/log/nginx/error.logPlain text.
PHP-FPM/var/log/php-fpm/php-fpm.logPool warnings, max_children hits.
PHP slow log/var/log/php-fpm/php-slow.logStack traces of slow requests, multiline (Hypernode docs).
MySQL slow log/var/log/mysql/mysql-slow.logThreshold set by the mysql_long_query_time setting. Multiline.
Magento<magento root>/var/log/*.logsystem, exception, and whatever your modules write. Multiline.

Two Hypernode details catch people out.

Log rotation uses copytruncate. Hypernode's hypernode-auto-logrotate copies the file and truncates the original in place, so the shipper must cope with a file that suddenly shrinks rather than being swapped out. Vector's file source handles this, but make sure your globs only match the live *.log files and never the rotated .1.gz copies, or you ship everything twice.

Varnish means some requests log twice. With Varnish on, a cache miss goes front-end nginx, then Varnish, then a backend nginx vhost on 127.0.0.1:8080, and both nginx hops write an access line. Don't drop either (the backend line tells you the request actually hit PHP), tag them instead. The JSON already carries a port field, so that is one line of config.

Also, if you deploy with release folders, point the Magento glob at the shared var/ directory, not the release symlink, otherwise every deploy looks like a brand new set of files.

The shipper: Vector

I picked Vector because it is one static binary, it does not need root, and it can transform events before they leave the box. Grab the musl build (no glibc surprises) and put it under /data/web, outside your deploy folders so a release prune never removes it:

VER=0.58.0
mkdir -p /data/web/vector/{bin,config,data,log}
curl -sSfL "https://packages.timber.io/vector/${VER}/vector-${VER}-x86_64-unknown-linux-musl.tar.gz" \
  | tar -xz -C /tmp
cp /tmp/vector-x86_64-unknown-linux-musl/bin/vector /data/web/vector/bin/
/data/web/vector/bin/vector --version

Here is a trimmed config showing the pattern. I have left out the error, FPM and MySQL sources since they are the same shape as the Magento one with a different path and start pattern.

/data/web/vector/config/vector.yaml

data_dir: /data/web/vector/data

sources:
  nginx_access:
    type: file
    include: [/var/log/nginx/access.log]
    read_from: end
  magento:
    type: file
    include: [/data/web/magento2/var/log/*.log]
    read_from: end
    multiline:                       # a stack trace stays one event
      mode: halt_before
      start_pattern: '^\[\d{4}-\d{2}-\d{2}'
      condition_pattern: '^\[\d{4}-\d{2}-\d{2}'
      timeout_ms: 1000

transforms:
  parse_nginx:
    type: remap
    inputs: [nginx_access]
    source: |
      .source = "nginx_access"
      .env = "${SHOP_ENV}"
      parsed, err = parse_json(.message)
      if err == null && is_object(parsed) {
        . = merge(., object!(parsed))
      }
      port = to_string(.port) ?? ""
      .hop = if port == "8080" { "backend" } else { "front" }
  parse_magento:
    type: remap
    inputs: [magento]
    source: |
      .source = "magento"
      .env = "${SHOP_ENV}"
  redact:
    type: remap
    inputs: [parse_nginx, parse_magento]
    drop_on_error: true              # never ship an event we failed to clean
    source: |
      .message = to_string(.message) ?? ""
      .message = truncate(.message, 65536)
      .message = redact(.message, filters: [
        r'[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}',
        r'\b(?:\d[ -]?){12,18}\d\b',
        r'(?i)(?:password|passwd|token|secret|key)=[^&\s"]+'
      ])

sinks:
  victorialogs:
    type: elasticsearch
    inputs: [redact]
    endpoints: ["https://logs.example.com/insert/elasticsearch/"]
    api_version: v8
    compression: gzip
    healthcheck:
      enabled: false
    query:
      _msg_field: message
      _time_field: timestamp
      _stream_fields: host,env,source
    request:
      headers:
        Authorization: "Bearer ${LOG_WRITE_TOKEN}"
    buffer:
      type: disk                     # survives the log box being down
      max_size: 268435488
      when_full: block

The sink settings come straight from VictoriaLogs' own Vector setup page, the only additions are the bearer token header and the disk buffer. The redaction filters are only an example, extend them to whatever your shop's URLs actually leak (the same redact() call is worth applying to the parsed request and referer fields too).

Redact on the node, not at the far end. Once an email address or a reset token has crossed the wire, it is in the database, the backups and whatever the agent quoted back to you. Cleaning it before it leaves is the only version of this that holds up.

Vector reads the token and environment name from a small env file, so the config itself carries no secrets. A wrapper loads it and hands over to Vector:

/data/web/vector/run-vector.sh

#!/bin/bash
set -a
. /data/web/vector/config/vector.env     # SHOP_ENV=prod, LOG_WRITE_TOKEN=...
set +a
exec /data/web/vector/bin/vector --config /data/web/vector/config/vector.yaml

And the Supervisor program that keeps it running:

/data/web/supervisor/vector.conf

[program:vector]
command=/data/web/vector/run-vector.sh
autostart=true
autorestart=true
redirect_stderr=true
stdout_logfile=/data/web/vector/log/vector.log

Keep Vector's own log out of every glob it tails (it lives under /data/web/vector/log above for that reason), otherwise it happily ships its own complaints about shipping.

I install all of this from a deploy step rather than by hand, an idempotent install.sh in the repo that the deploy pipeline runs after each release. Two rules for that script. It always exits 0, a broken log shipper must never fail a shop deploy. And it has a kill switch, if logging is disabled for that environment or the token is missing it removes the program from Supervisor and stops, rather than leaving something half configured running.

The receiving end: VictoriaLogs behind Caddy

I looked at an ELK stack for about five minutes, it is far too much machine for a handful of shops. VictoriaLogs is a single binary, runs happily in a few hundred MB of RAM, and has proper multi-tenancy. Mine runs in Docker on a small VPS with retention and a disk cap set, so a bot flood can't eat the disk:

victorialogs:
  image: victoriametrics/victoria-logs:v1.52.0
  command:
    - -storageDataPath=/vlogs
    - -retentionPeriod=30d
    - -retention.maxDiskSpaceUsageBytes=20GiB
    - -httpListenAddr=:9428
  mem_limit: 600m
  expose: ["9428"]                 # never published, only Caddy can reach it

Now the important part. VictoriaLogs picks the tenant from AccountID and ProjectID request headers (multitenancy docs). If clients could set those themselves, any token would read any shop. So Caddy maps each token to exactly one tenant and overwrites those headers, whatever the client sent:

Caddyfile

logs.example.com {
	@shop_prod_write {
		path /insert/*
		header Authorization "Bearer {$SHOP_PROD_WRITE}"
	}
	@shop_prod_read {
		path /select/logsql/*
		header Authorization "Bearer {$SHOP_PROD_READ}"
	}
	handle @shop_prod_write {
		request_header AccountID 12
		request_header ProjectID 34
		request_header -Authorization
		reverse_proxy victorialogs:9428
	}
	handle @shop_prod_read {
		request_header AccountID 12
		request_header ProjectID 34
		request_header -Authorization
		reverse_proxy victorialogs:9428
	}
	handle {
		respond 403                # everything else, including the web UI and admin endpoints
	}
}

Repeat the pair per shop and environment. That gives you:

The agent end: mcp-victorialogs

VictoriaMetrics publish an official MCP server, mcp-victorialogs, which wraps the read-only query APIs as tools (query, list streams and fields, hit stats) and carries its own copy of the docs. I run one instance on my dev machine in HTTP mode, and let it pass the caller's Authorization header straight through to Caddy:

docker run -d --name mcp-victorialogs --restart always \
  -e VL_INSTANCE_ENTRYPOINT=https://logs.example.com \
  -e MCP_SERVER_MODE=http \
  -e MCP_LISTEN_ADDR=:8081 \
  -e MCP_PASSTHROUGH_HEADERS=Authorization \
  -p 127.0.0.1:8081:8081 \
  ghcr.io/victoriametrics/mcp-victorialogs:v1.9.0

The MCP server holds no tokens and no data. Each project's agent config brings its own read token, so one shared server still keeps every shop separated:

.mcp.json (per project)

"logs-prod": {
  "type": "http",
  "url": "http://localhost:8081/mcp",
  "headers": { "Authorization": "Bearer ${LOGS_READ_TOKEN_PROD}" }
}

After that, "what were the top 5xx URLs on prod in the last hour?" turns into a LogsQL query the agent writes itself:

_time:1h source:nginx_access status:>=500
  | stats by (request) count() hits
  | sort by (hits desc)
  | limit 10

A tip from my own setup: put a decent description on the MCP entry saying what the sources and stream fields are, and that the agent should ask the error tracker first for stack traces. Agents pick tools off that text, a vague description gets you a vague query.

Things that bit me

HTTP 200, zero rows. When I tested ingestion by hand with curl against the /insert/jsonline endpoint, it returned 200 and stored nothing. Without Content-Type: application/stream+json the body got parsed as a form and quietly thrown away. Always check with a count query after a test insert, a 200 alone proves nothing.

Where this is at

The receiving side, the tenancy rules and the MCP wiring are live and tested, including the forged-header checks above. The shop side (Vector config, redaction tests, install script, Supervisor program) is going through my normal plan, test and uat-verify cycle as I write this, so treat the Vector config here as a solid starting point rather than a finished, battle-tested one. I will update this post with real numbers (memory use on the node, lag from write to query) once it has run on uat for a while.

Everything this post leans on, the Hypernode docs first.

I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)

More in this series: The Hour My Agents Spent Waiting For Nothing — why my agents are not allowed to background work they need the result of. · Smaller, Then Green — a mandatory cut-then-test loop that trims over-built code out of every plan's diff before any reviewer sees it. · Plans Run By A Fleet — how features run through a planned, adversarially-reviewed, parallel TDD pipeline instead of one long agent chat. · Hooks, Not Hopes — sorting every agent rule into prose or a blocking hook, and proving the hooks still fire. · The Agent Chatroom — giving the agents across my fleet a threaded chatroom so I stopped being the message bus. · The Commit That Has To Prove Itself — gating every commit behind recorded, state-hashed test evidence. · My AI Is Not Allowed To Guess — forcing blast-radius-first investigation and banning blame-shift excuses. · The Near-Miss That Banned Summaries — banning summarised page-fetches fleet-wide after a near-miss on a security advisory. · The Code Quality Checks Nobody Runs — coding standards, static analysis and comment hygiene moved into the commit path. · The ProxiBlue Debugger Discipline — blocking var_dump and wiring in real breakpoint debugging. · The ProxiBlue Domain Graph — giving my AI coding agents long-term domain knowledge with a temporal knowledge graph. · The ProxiBlue Falsifiable Rulebook — testing the rules and guard hooks that govern my AI coding agents.