dpm_bush1. Linux production mistakes that are easy to miss A production incident doesn’t always...
A production incident doesn’t always start with a dramatic crash. Sometimes a deploy works under your account but fails for the service user, a disk fills because deleted logs are still open, or a configuration edit takes effect only after an unexpected restart.
A command that succeeds in your shell may fail in a service. Your login account might have different group membership, environment variables, working directory, or access to a secret file than the account running the application.
Check the configured identity and paths instead of assuming they match your shell:
systemctl show myapp -p User -p Group -p WorkingDirectory
If practical, test a read or write as that account using sudo -u appuser .... For file access, inspect every parent directory, not just the file itself: a process needs search (x) permission on each directory in the path.
Variables exported in .bashrc or an interactive shell generally aren’t available to a system service. Put service-specific settings in a deliberate place, such as a systemd unit’s EnvironmentFile=, and restrict that file’s access if it contains secrets. After editing a unit, reload systemd’s unit definitions and restart or reload the service as appropriate.
A filesystem can have capacity remaining but no inodes available for new files. Check both:
df -h /
df -i /
A workload that creates many tiny cache files can exhaust inodes before it uses much disk space. Identify the directory generating files and fix its retention behavior rather than deleting random data.
Removing a large log file doesn’t necessarily return its space while a running process still holds it open. If df reports a full filesystem but directory totals don’t account for it, inspect open deleted files:
sudo lsof +L1
Use the application’s supported log-reopen mechanism or a planned service restart. Avoid truncating arbitrary open files without understanding the application’s logging behavior.
Before changing a production config, keep a known-good copy, validate syntax with the application’s own checker, and apply the change in a way you can reverse. For SSH or firewall changes, keep an existing session open and verify a second connection before closing it. For systemd unit edits, systemctl cat myapp shows the effective unit and its drop-ins.
Prevention is mostly about making assumptions visible: which user runs the service, which filesystem stores its data, and which exact config file the process reads. For interpreting service state before changing anything, the systemctl status guide is a useful reference.
When you run ssh user@server, the client first establishes a network connection to the server’s SSH daemon. Before sending a password or accepting a public-key login, the two sides need to agree on how to protect the connection and establish that they are talking to the intended server.
The server presents a host key. Your client compares it with the host key saved for that host in known_hosts. A first connection may ask you to trust a new key; a changed key produces a warning because the server’s identity no longer matches what you previously recorded.
That warning is not proof of an attack, but it is a reason to verify the change through a trusted channel before updating your stored record. Host verification answers: “Is this the server I expected?” It does not prove that your user account is authorized.
Client and server negotiate supported algorithms, then perform a key exchange. The exchange lets both sides derive shared session keys without sending those keys directly over the network. The host key authenticates the server during this process; it is not normally the key used to encrypt every packet.
After the exchange, SSH encrypts and integrity-protects the transport. The client and server can then send authentication messages inside that protected channel.
Only after transport protection is established does the client authenticate the requested account. With public-key authentication, the client proves possession of a private key by signing data for the session. The server checks that proof against public keys authorized for the account. Password authentication, when enabled, is also carried through the encrypted transport.
So there are two distinct checks: the client verifies the server’s host identity, and the server verifies the user’s credentials. A successful user login does not make it safe to ignore a host-key warning.
For the practical side of checking a first connection or a changed fingerprint, see the host-key verification guide.
A low-cost VPS can run useful services, but it has little room for unlimited growth. The most reliable design is usually the one with the fewest moving parts and a clear limit for each resource.
Separate irreplaceable state from rebuildable software. Database files, uploaded media, and configuration belong in a backup plan. Containers and packages can often be recreated, but only if you have recorded how they were configured and can retrieve the needed images or packages.
Keep at least one backup somewhere other than the VPS. A backup on the same disk won’t help if the instance or its storage disappears. Periodically test restoring a small sample; a backup you have never restored is an assumption, not a verified recovery process.
A database, reverse proxy, application, and monitoring stack all compete for memory. Leave headroom for package upgrades, traffic spikes, and filesystem cache. If the VPS starts swapping heavily or killing processes under memory pressure, adding more services may make reliability worse. Check current pressure with tools such as free -h and vmstat, and inspect kernel messages if processes disappear unexpectedly.
Set retention for logs, backups, and uploaded files. Estimate growth instead of waiting for an alert at 100% disk usage. On a small disk, a few retained database dumps can matter as much as application data.
Expose only services that need inbound access. Bind databases and internal dashboards to loopback or a private network where possible; allow public access only where required by the application. Check both the VPS firewall and any provider-level firewall because either can block or permit traffic independently.
A monitoring check should answer a useful question: can the public endpoint respond, is the disk approaching a limit, and is the backup job recent? A simple external availability check catches failures that local service status cannot, such as a broken firewall rule or DNS issue.
Keep services under a manager such as systemd or a container orchestrator rather than starting them in an SSH shell. Document the commands and data paths needed to rebuild the host. On a small VPS, boring operations are a feature.
When something breaks, resist changing several things at once. First describe the symptom precisely, then find the earliest layer where expected behavior stops.
Write down what failed, when it began, and what changed shortly beforehand. Does the failure affect one user, one host, one service, or every service? Is it continuous or intermittent? A timestamp helps correlate an application error with system logs and deployments.
For a web application, the path might be DNS → network connection → reverse proxy → upstream service → database. Test each boundary separately. A successful DNS lookup doesn’t prove the server is reachable; a listening socket doesn’t prove the application can serve a request.
Inspect the relevant service’s recent logs around the failure time:
journalctl -u myapp --since "20 minutes ago"
Then check whether the process exists and what resources it is using:
ps -o pid,ppid,stat,%cpu,%mem,cmd -C myapp
The process name may differ from the service name, so use systemctl status myapp to confirm the unit and its main process. If it is running, check whether it is listening on the address and port the next layer expects. A service bound only to 127.0.0.1 is not reachable through its private network address.
Suppose a deployment returns gateway errors. A useful sequence is: confirm the proxy received the request, inspect its upstream error, check whether the application service is active, test the upstream locally, and then examine application logs. If the upstream refuses a connection, investigate its listener and startup state before changing proxy timeouts.
Make one reversible change, repeat the same test, and record the result. If the result does not fit the hypothesis, undo the change and move to the next boundary. This avoids accumulating unrelated edits that make the original cause harder to find.
For log filtering and time-based inspection, the journalctl guide covers useful query patterns. For sockets, combine local listener checks with a test from the client’s network; each answers a different question.
The fifth direction was cut off after “How to manage,” so this section takes a useful adjacent angle: managing recurring server maintenance without turning it into a risky batch of unreviewed changes.
Start by checking available updates and reading what they will change. Apply updates in a planned window when a reboot or service restart could affect users. For critical hosts, know how to access the provider’s console before changing networking, boot, or authentication settings.
For each meaningful change, record the date, purpose, files or services affected, and rollback step. This can be a ticket, a small operations log, or version-controlled configuration. The goal is not paperwork; it is being able to answer “what changed?” during an incident.
After an update or configuration change, check the affected service, test its real endpoint, and confirm scheduled jobs still run. Review disk usage and backup freshness on a recurring schedule. Automate checks where possible, but make failures visible rather than silently discarding their output.
A server is easier to maintain when routine work is small, repeatable, and reversible. Treat updates, backups, and monitoring as one operational loop: prepare, change, verify, and retain a way back.
I'm building SSHFlow — an SSH client that organizes terminals, SFTP, code editing, and server tools into workspaces.