Engineering Notes · agent operations · observability
My AI agents are not allowed to SSH to a live shop. Fine, except the answer to half the questions I ask them is sitting in an nginx or PHP-FPM log on that box. So the logs come to the agent instead. Hypernode's Supervisor feature runs a small shipper on the node, a self-hosted VictoriaLogs instance holds the data, and an MCP server gives the agent read-only, per-shop search.
I already ship exceptions to Bugsink, and that covers the stack traces. It does not cover the rest of what you look at when a shop misbehaves: the burst of 502s at 3am, the one category page taking 12 seconds, the slow query log filling up after a reindex, the bot hammering the search endpoint. That all lives in plain log files on the Hypernode.
Reading those files used to mean someone SSHing in. I don't want agents doing that on live (read-only is a promise, not a guarantee, and one stray command is all it takes), and I don't want to be the human copy-paste layer between the server and the agent either. So the goal here is simple: get the logs off the box, keep them per shop and per environment, and let the agent query them through a tool that physically cannot write anything.
Everything on the shop side is push-only. Nothing connects in to the Hypernode, and nothing on the node needs root.
The piece that makes this possible is a Hypernode feature a lot of people never turn on. Hypernode ships Supervisor for long running programs, it watches whatever you give it and restarts it if it dies. It is off by default, and you turn it on with the hypernode-systemctl tool:
hypernode-systemctl settings supervisor_enabled
hypernode-systemctl settings supervisor_enabled True
(The first line just tells you the current value, worth running before you change anything.) If you deploy with Hypernode Deploy, you can keep this in source control instead, it has a HypernodeSettingConfiguration and a SupervisorConfiguration that set the flag and install your program configs on every deploy.
Programs are declared as *.conf files in /data/web/supervisor/, then loaded with supervisorctl:
supervisorctl reread
supervisorctl add vector
supervisorctl status
Supervisor's own logs end up in /data/var/log/supervisor/, which is the first place to look if something will not stay up.
Hypernode is quite good here, most of the useful logs are already readable by the app user and some are already structured. These are the ones I ship (check them on your own node first, plans and setups differ):
| Log | Path | Notes |
|---|---|---|
| nginx access | /var/log/nginx/access.log | Already JSON, one object per line (Hypernode docs). No grok patterns needed. |
| nginx error | /var/log/nginx/error.log | Plain text. |
| PHP-FPM | /var/log/php-fpm/php-fpm.log | Pool warnings, max_children hits. |
| PHP slow log | /var/log/php-fpm/php-slow.log | Stack traces of slow requests, multiline (Hypernode docs). |
| MySQL slow log | /var/log/mysql/mysql-slow.log | Threshold set by the mysql_long_query_time setting. Multiline. |
| Magento | <magento root>/var/log/*.log | system, exception, and whatever your modules write. Multiline. |
Two Hypernode details catch people out.
Log rotation uses copytruncate. Hypernode's hypernode-auto-logrotate copies the file and truncates the original in place, so the shipper must cope with a file that suddenly shrinks rather than being swapped out. Vector's file source handles this, but make sure your globs only match the live *.log files and never the rotated .1.gz copies, or you ship everything twice.
Varnish means some requests log twice. With Varnish on, a cache miss goes front-end nginx, then Varnish, then a backend nginx vhost on 127.0.0.1:8080, and both nginx hops write an access line. Don't drop either (the backend line tells you the request actually hit PHP), tag them instead. The JSON already carries a port field, so that is one line of config.
Also, if you deploy with release folders, point the Magento glob at the shared var/ directory, not the release symlink, otherwise every deploy looks like a brand new set of files.
I picked Vector because it is one static binary, it does not need root, and it can transform events before they leave the box. Grab the musl build (no glibc surprises) and put it under /data/web, outside your deploy folders so a release prune never removes it:
VER=0.58.0
mkdir -p /data/web/vector/{bin,config,data,log}
curl -sSfL "https://packages.timber.io/vector/${VER}/vector-${VER}-x86_64-unknown-linux-musl.tar.gz" \
| tar -xz -C /tmp
cp /tmp/vector-x86_64-unknown-linux-musl/bin/vector /data/web/vector/bin/
/data/web/vector/bin/vector --version
Here is a trimmed config showing the pattern. I have left out the error, FPM and MySQL sources since they are the same shape as the Magento one with a different path and start pattern.
/data/web/vector/config/vector.yaml
data_dir: /data/web/vector/data
sources:
nginx_access:
type: file
include: [/var/log/nginx/access.log]
read_from: end
magento:
type: file
include: [/data/web/magento2/var/log/*.log]
read_from: end
multiline: # a stack trace stays one event
mode: halt_before
start_pattern: '^\[\d{4}-\d{2}-\d{2}'
condition_pattern: '^\[\d{4}-\d{2}-\d{2}'
timeout_ms: 1000
transforms:
parse_nginx:
type: remap
inputs: [nginx_access]
source: |
.source = "nginx_access"
.env = "${SHOP_ENV}"
parsed, err = parse_json(.message)
if err == null && is_object(parsed) {
. = merge(., object!(parsed))
}
port = to_string(.port) ?? ""
.hop = if port == "8080" { "backend" } else { "front" }
parse_magento:
type: remap
inputs: [magento]
source: |
.source = "magento"
.env = "${SHOP_ENV}"
redact:
type: remap
inputs: [parse_nginx, parse_magento]
drop_on_error: true # never ship an event we failed to clean
source: |
.message = to_string(.message) ?? ""
.message = truncate(.message, 65536)
.message = redact(.message, filters: [
r'[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}',
r'\b(?:\d[ -]?){12,18}\d\b',
r'(?i)(?:password|passwd|token|secret|key)=[^&\s"]+'
])
sinks:
victorialogs:
type: elasticsearch
inputs: [redact]
endpoints: ["https://logs.example.com/insert/elasticsearch/"]
api_version: v8
compression: gzip
healthcheck:
enabled: false
query:
_msg_field: message
_time_field: timestamp
_stream_fields: host,env,source
request:
headers:
Authorization: "Bearer ${LOG_WRITE_TOKEN}"
buffer:
type: disk # survives the log box being down
max_size: 268435488
when_full: block
The sink settings come straight from VictoriaLogs' own Vector setup page, the only additions are the bearer token header and the disk buffer. The redaction filters are only an example, extend them to whatever your shop's URLs actually leak (the same redact() call is worth applying to the parsed request and referer fields too).
Redact on the node, not at the far end. Once an email address or a reset token has crossed the wire, it is in the database, the backups and whatever the agent quoted back to you. Cleaning it before it leaves is the only version of this that holds up.
Vector reads the token and environment name from a small env file, so the config itself carries no secrets. A wrapper loads it and hands over to Vector:
/data/web/vector/run-vector.sh
#!/bin/bash
set -a
. /data/web/vector/config/vector.env # SHOP_ENV=prod, LOG_WRITE_TOKEN=...
set +a
exec /data/web/vector/bin/vector --config /data/web/vector/config/vector.yaml
And the Supervisor program that keeps it running:
/data/web/supervisor/vector.conf
[program:vector]
command=/data/web/vector/run-vector.sh
autostart=true
autorestart=true
redirect_stderr=true
stdout_logfile=/data/web/vector/log/vector.log
Keep Vector's own log out of every glob it tails (it lives under /data/web/vector/log above for that reason), otherwise it happily ships its own complaints about shipping.
I install all of this from a deploy step rather than by hand, an idempotent install.sh in the repo that the deploy pipeline runs after each release. Two rules for that script. It always exits 0, a broken log shipper must never fail a shop deploy. And it has a kill switch, if logging is disabled for that environment or the token is missing it removes the program from Supervisor and stops, rather than leaving something half configured running.
I looked at an ELK stack for about five minutes, it is far too much machine for a handful of shops. VictoriaLogs is a single binary, runs happily in a few hundred MB of RAM, and has proper multi-tenancy. Mine runs in Docker on a small VPS with retention and a disk cap set, so a bot flood can't eat the disk:
victorialogs:
image: victoriametrics/victoria-logs:v1.52.0
command:
- -storageDataPath=/vlogs
- -retentionPeriod=30d
- -retention.maxDiskSpaceUsageBytes=20GiB
- -httpListenAddr=:9428
mem_limit: 600m
expose: ["9428"] # never published, only Caddy can reach it
Now the important part. VictoriaLogs picks the tenant from AccountID and ProjectID request headers (multitenancy docs). If clients could set those themselves, any token would read any shop. So Caddy maps each token to exactly one tenant and overwrites those headers, whatever the client sent:
Caddyfile
logs.example.com {
@shop_prod_write {
path /insert/*
header Authorization "Bearer {$SHOP_PROD_WRITE}"
}
@shop_prod_read {
path /select/logsql/*
header Authorization "Bearer {$SHOP_PROD_READ}"
}
handle @shop_prod_write {
request_header AccountID 12
request_header ProjectID 34
request_header -Authorization
reverse_proxy victorialogs:9428
}
handle @shop_prod_read {
request_header AccountID 12
request_header ProjectID 34
request_header -Authorization
reverse_proxy victorialogs:9428
}
handle {
respond 403 # everything else, including the web UI and admin endpoints
}
}
Repeat the pair per shop and environment. That gives you:
/insert/*, so a token sitting on a shop server can push logs but never read them.AccountID/ProjectID headers does nothing, Caddy overwrites them. I tested this by sending a uat token with forged prod headers, and got uat data back.VictoriaMetrics publish an official MCP server, mcp-victorialogs, which wraps the read-only query APIs as tools (query, list streams and fields, hit stats) and carries its own copy of the docs. I run one instance on my dev machine in HTTP mode, and let it pass the caller's Authorization header straight through to Caddy:
docker run -d --name mcp-victorialogs --restart always \
-e VL_INSTANCE_ENTRYPOINT=https://logs.example.com \
-e MCP_SERVER_MODE=http \
-e MCP_LISTEN_ADDR=:8081 \
-e MCP_PASSTHROUGH_HEADERS=Authorization \
-p 127.0.0.1:8081:8081 \
ghcr.io/victoriametrics/mcp-victorialogs:v1.9.0
The MCP server holds no tokens and no data. Each project's agent config brings its own read token, so one shared server still keeps every shop separated:
.mcp.json (per project)
"logs-prod": {
"type": "http",
"url": "http://localhost:8081/mcp",
"headers": { "Authorization": "Bearer ${LOGS_READ_TOKEN_PROD}" }
}
After that, "what were the top 5xx URLs on prod in the last hour?" turns into a LogsQL query the agent writes itself:
_time:1h source:nginx_access status:>=500
| stats by (request) count() hits
| sort by (hits desc)
| limit 10
A tip from my own setup: put a decent description on the MCP entry saying what the sources and stream fields are, and that the agent should ask the error tracker first for stack traces. Agents pick tools off that text, a vague description gets you a vague query.
HTTP 200, zero rows. When I tested ingestion by hand with curl against the /insert/jsonline endpoint, it returned 200 and stored nothing. Without Content-Type: application/stream+json the body got parsed as a form and quietly thrown away. Always check with a count query after a test insert, a 200 alone proves nothing.
The receiving side, the tenancy rules and the MCP wiring are live and tested, including the forged-header checks above. The shop side (Vector config, redaction tests, install script, Supervisor program) is going through my normal plan, test and uat-verify cycle as I write this, so treat the Vector config here as a solid starting point rather than a finished, battle-tested one. I will update this post with real numbers (memory use on the node, lag from write to query) once it has run on uat for a while.
Everything this post leans on, the Hypernode docs first.
I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)