Five Linux Field Guides: Production Traps, SSH, Small VPSes, Debugging, and Maintenance

# programming# productivity# tutorial# devops
Five Linux Field Guides: Production Traps, SSH, Small VPSes, Debugging, and Maintenancedpm_bush

1. Linux production mistakes that are easy to miss A production incident doesn’t always...

1. Linux production mistakes that are easy to miss

A production incident doesn’t always start with a dramatic crash. Sometimes a deploy works under your account but fails for the service user, a disk fills because deleted logs are still open, or a configuration edit takes effect only after an unexpected restart.

Testing as the wrong user

A command that succeeds in your shell may fail in a service. Your login account might have different group membership, environment variables, working directory, or access to a secret file than the account running the application.

Check the configured identity and paths instead of assuming they match your shell:

systemctl show myapp -p User -p Group -p WorkingDirectory
Enter fullscreen mode Exit fullscreen mode

If practical, test a read or write as that account using sudo -u appuser .... For file access, inspect every parent directory, not just the file itself: a process needs search (x) permission on each directory in the path.

Assuming a service inherited your shell environment

Variables exported in .bashrc or an interactive shell generally aren’t available to a system service. Put service-specific settings in a deliberate place, such as a systemd unit’s EnvironmentFile=, and restrict that file’s access if it contains secrets. After editing a unit, reload systemd’s unit definitions and restart or reload the service as appropriate.

Confusing free space with free inodes

A filesystem can have capacity remaining but no inodes available for new files. Check both:

df -h /
df -i /
Enter fullscreen mode Exit fullscreen mode

A workload that creates many tiny cache files can exhaust inodes before it uses much disk space. Identify the directory generating files and fix its retention behavior rather than deleting random data.

Deleting a log that a process still has open

Removing a large log file doesn’t necessarily return its space while a running process still holds it open. If df reports a full filesystem but directory totals don’t account for it, inspect open deleted files:

sudo lsof +L1
Enter fullscreen mode Exit fullscreen mode

Use the application’s supported log-reopen mechanism or a planned service restart. Avoid truncating arbitrary open files without understanding the application’s logging behavior.

Changing configuration without a rollback path

Before changing a production config, keep a known-good copy, validate syntax with the application’s own checker, and apply the change in a way you can reverse. For SSH or firewall changes, keep an existing session open and verify a second connection before closing it. For systemd unit edits, systemctl cat myapp shows the effective unit and its drop-ins.

Prevention is mostly about making assumptions visible: which user runs the service, which filesystem stores its data, and which exact config file the process reads. For interpreting service state before changing anything, the systemctl status guide is a useful reference.


2. SSH is a sequence of trust decisions, not just a login prompt

When you run ssh user@server, the client first establishes a network connection to the server’s SSH daemon. Before sending a password or accepting a public-key login, the two sides need to agree on how to protect the connection and establish that they are talking to the intended server.

The server’s identity comes first

The server presents a host key. Your client compares it with the host key saved for that host in known_hosts. A first connection may ask you to trust a new key; a changed key produces a warning because the server’s identity no longer matches what you previously recorded.

That warning is not proof of an attack, but it is a reason to verify the change through a trusted channel before updating your stored record. Host verification answers: “Is this the server I expected?” It does not prove that your user account is authorized.

Key exchange establishes protected transport

Client and server negotiate supported algorithms, then perform a key exchange. The exchange lets both sides derive shared session keys without sending those keys directly over the network. The host key authenticates the server during this process; it is not normally the key used to encrypt every packet.

After the exchange, SSH encrypts and integrity-protects the transport. The client and server can then send authentication messages inside that protected channel.

User authentication is a separate question

Only after transport protection is established does the client authenticate the requested account. With public-key authentication, the client proves possession of a private key by signing data for the session. The server checks that proof against public keys authorized for the account. Password authentication, when enabled, is also carried through the encrypted transport.

So there are two distinct checks: the client verifies the server’s host identity, and the server verifies the user’s credentials. A successful user login does not make it safe to ignore a host-key warning.

For the practical side of checking a first connection or a changed fingerprint, see the host-key verification guide.


3. A small VPS needs budgets more than clever architecture

A low-cost VPS can run useful services, but it has little room for unlimited growth. The most reliable design is usually the one with the fewest moving parts and a clear limit for each resource.

Decide what must survive

Separate irreplaceable state from rebuildable software. Database files, uploaded media, and configuration belong in a backup plan. Containers and packages can often be recreated, but only if you have recorded how they were configured and can retrieve the needed images or packages.

Keep at least one backup somewhere other than the VPS. A backup on the same disk won’t help if the instance or its storage disappears. Periodically test restoring a small sample; a backup you have never restored is an assumption, not a verified recovery process.

Budget memory and disk

A database, reverse proxy, application, and monitoring stack all compete for memory. Leave headroom for package upgrades, traffic spikes, and filesystem cache. If the VPS starts swapping heavily or killing processes under memory pressure, adding more services may make reliability worse. Check current pressure with tools such as free -h and vmstat, and inspect kernel messages if processes disappear unexpectedly.

Set retention for logs, backups, and uploaded files. Estimate growth instead of waiting for an alert at 100% disk usage. On a small disk, a few retained database dumps can matter as much as application data.

Keep the public surface small

Expose only services that need inbound access. Bind databases and internal dashboards to loopback or a private network where possible; allow public access only where required by the application. Check both the VPS firewall and any provider-level firewall because either can block or permit traffic independently.

Make recovery observable

A monitoring check should answer a useful question: can the public endpoint respond, is the disk approaching a limit, and is the backup job recent? A simple external availability check catches failures that local service status cannot, such as a broken firewall rule or DNS issue.

Keep services under a manager such as systemd or a container orchestrator rather than starting them in an SSH shell. Document the commands and data paths needed to rebuild the host. On a small VPS, boring operations are a feature.


4. Troubleshoot Linux by narrowing the failure boundary

When something breaks, resist changing several things at once. First describe the symptom precisely, then find the earliest layer where expected behavior stops.

Capture the scope and time

Write down what failed, when it began, and what changed shortly beforehand. Does the failure affect one user, one host, one service, or every service? Is it continuous or intermittent? A timestamp helps correlate an application error with system logs and deployments.

Follow the request path

For a web application, the path might be DNS → network connection → reverse proxy → upstream service → database. Test each boundary separately. A successful DNS lookup doesn’t prove the server is reachable; a listening socket doesn’t prove the application can serve a request.

Inspect the relevant service’s recent logs around the failure time:

journalctl -u myapp --since "20 minutes ago"
Enter fullscreen mode Exit fullscreen mode

Then check whether the process exists and what resources it is using:

ps -o pid,ppid,stat,%cpu,%mem,cmd -C myapp
Enter fullscreen mode Exit fullscreen mode

The process name may differ from the service name, so use systemctl status myapp to confirm the unit and its main process. If it is running, check whether it is listening on the address and port the next layer expects. A service bound only to 127.0.0.1 is not reachable through its private network address.

Use one hypothesis per change

Suppose a deployment returns gateway errors. A useful sequence is: confirm the proxy received the request, inspect its upstream error, check whether the application service is active, test the upstream locally, and then examine application logs. If the upstream refuses a connection, investigate its listener and startup state before changing proxy timeouts.

Make one reversible change, repeat the same test, and record the result. If the result does not fit the hypothesis, undo the change and move to the next boundary. This avoids accumulating unrelated edits that make the original cause harder to find.

For log filtering and time-based inspection, the journalctl guide covers useful query patterns. For sockets, combine local listener checks with a test from the client’s network; each answers a different question.


5. A practical maintenance routine for a Linux server

The fifth direction was cut off after “How to manage,” so this section takes a useful adjacent angle: managing recurring server maintenance without turning it into a risky batch of unreviewed changes.

Separate inspection from action

Start by checking available updates and reading what they will change. Apply updates in a planned window when a reboot or service restart could affect users. For critical hosts, know how to access the provider’s console before changing networking, boot, or authentication settings.

Keep a short change record

For each meaningful change, record the date, purpose, files or services affected, and rollback step. This can be a ticket, a small operations log, or version-controlled configuration. The goal is not paperwork; it is being able to answer “what changed?” during an incident.

Verify after maintenance

After an update or configuration change, check the affected service, test its real endpoint, and confirm scheduled jobs still run. Review disk usage and backup freshness on a recurring schedule. Automate checks where possible, but make failures visible rather than silently discarding their output.

A server is easier to maintain when routine work is small, repeatable, and reversible. Treat updates, backups, and monitoring as one operational loop: prepare, change, verify, and retain a way back.

I'm building SSHFlow — an SSH client that organizes terminals, SFTP, code editing, and server tools into workspaces.