Chapter 2 — The Linux Operating System
Overview
Nearly every server I've had to patch or dig through logs on — web server, database, container — runs Linux, so whatever traces an attack leaves behind live in the mechanisms of that same operating system. This chapter looks at Linux through that lens: not to administer it, but to know who just did what, under which privileges, and why they were allowed to.
The starting point is the boundary every Unix system is built on: the kernel runs at ring 0 with exclusive access to hardware, while application processes are confined to user space and must go through a syscall for every resource they need — a barrier that isolates faults and malware from both the hardware and other processes. On top of that, Linux organizes files under the FHS standard (configuration in /etc, logs in /var/log) so tools and monitoring rules work the same regardless of distribution, then controls who can do what through the rwx permission model plus the special SUID/SGID/sticky bits — a misconfigured SUID-root binary is still a textbook privilege-escalation path. User identity lives in /etc/passwd, while passwords exist only as hashes in /etc/shadow, readable by root alone; the actual gate for running commands with elevated rights is sudo and PAM, which together give both least privilege and an audit trail.
The next layer is the process lifecycle: a process is born through fork() and then loads a new program image via execve(), carrying its own security identity — and during an investigation, a web server that suddenly spawns a shell is a textbook indicator of compromise. /proc exposes kernel state per process, making it a goldmine for forensics, while namespaces and cgroups are the two mechanisms — isolation and resource limiting — that every Docker container is built on. The full lifecycle of services on a machine, from startup through monitoring and automatic recovery, is handled by systemd via units, which is also what lets you harden individual services through declarative directives.
The rest of the chapter covers the day-to-day operations-and-investigation layer, which I've grouped into a short list since each piece stands fairly independently:
- Logging through
rsyslog/journaldcaptures every sensitive event, whilelogrotatekeeps the disk from filling up — but logs only earn their forensic value once they're shipped off the box before an attacker can wipe the local copy. - Package managers (
apt,dnf) don't just install software — they verify digital signatures in transit, which is the whole basis for trusting where anything running on the machine came from. - cron is as convenient for automation as it is attractive to attackers: a stray crontab entry is a quiet place to hide a backdoor that fires on schedule, which is why it's always on the list during an investigation.
- Under the hood of every command line is bash's file descriptors, redirection, and pipes chained into pipelines, plus text tools like
grep/awk/sed/sort/uniq— the toolkit for querying millions of log lines in a single command. - Last comes hardening: locking down SSH,
fail2banauto-blocking IPs that guess passwords, a firewall opening only the ports you need, and SELinux/AppArmor constraining a service's behavior even after it's already been compromised.
Commands and sample output throughout the chapter were actually run on Debian/Ubuntu or RHEL/Rocky; anything that depends on kernel or distro version is called out separately.
2.1. Overall architecture and the user space / kernel space boundary
Before diving into each mechanism, it is essential to understand the layered model, because nearly every security decision on Linux revolves around the boundary between user space and kernel space.
+---------------------------------------------------------------+
| USER SPACE (ring 3 on x86-64) |
| bash, sshd, nginx, python ... |
| libraries: glibc (libc.so.6), libssl ... |
| | system call (syscall instruction) |
+--------|------------------------------------------------------+
v
+---------------------------------------------------------------+
| KERNEL SPACE (ring 0) |
| - syscall dispatcher (sys_call_table) |
| - scheduler (CFS/EEVDF depending on version) |
| - VFS -> ext4/xfs/btrfs |
| - net stack (socket -> TCP/IP -> netfilter) |
| - LSM hooks (SELinux/AppArmor) |
| - process, memory (page table), namespaces, cgroups mgmt |
+---------------------------------------------------------------+
Why the ring 3 / ring 0 split? The x86-64 CPU has four privilege levels (rings 0–3); Linux uses only ring 0 (kernel) and ring 3 (user). They are separated so that user-space code cannot directly access the hardware, another process's page tables, or kernel structures — every request must go through a syscall, where the kernel checks permissions. This is the root security barrier of the whole system.
Syscalls at the instruction level (x86-64, Linux):
| Register | Role during a syscall |
|---|---|
rax |
Syscall number (e.g., read=0, write=1, open=2, execve=59) |
rdi |
Argument 1 |
rsi |
Argument 2 |
rdx |
Argument 3 |
r10 |
Argument 4 (note: NOT rcx as in the ordinary function-call ABI) |
r8 |
Argument 5 |
r9 |
Argument 6 |
rax (after the instruction) |
Return value (negative = -errno) |
The syscall instruction switches the CPU to ring 0 and jumps to the address in the LSTAR MSR. Observing a process's syscall sequence is a critical investigative technique:
strace -f -e trace=openat,connect,execve -s 200 curl -s https://example.com -o /dev/null
-f follows child processes too, -e trace= filters a group of syscalls, -s 200 prints up to 200 characters of strings. Sample output:
execve("/usr/bin/curl", ["curl","-s","https://example.com",...], 0x7ffd...) = 0
openat(AT_FDCWD, "/etc/ssl/certs/ca-certificates.crt", O_RDONLY) = 5
connect(6, {sa_family=AF_INET, sin_port=htons(443), sin_addr=inet_addr("192.0.2.10")}, 16) = -1 EINPROGRESS (Operation now in progress)
Security note: seccomp-bpf (used by container runtimes and by systemd's SystemCallFilter=) filters precisely these rax numbers. Understanding the syscall table helps in writing/reading seccomp profiles and detecting anomalous behavior (e.g., a web server suddenly calling execve).
2.2. Filesystem Hierarchy Standard (FHS)
The FHS (current standard, 3.0) defines the meaning of each top-level directory. Why have a standard? So that tools, scripts, and admins can be certain where files reside without depending on the distribution — essential when writing monitoring rules (e.g., tracking writes to /etc, /bin).
| Path | Purpose | Writable at runtime? | Security concern |
|---|---|---|---|
/ |
Root | — | — |
/bin, /sbin |
Essential binaries (often symlinked into /usr/bin) |
Should not be | Writing here = replacing a system binary |
/usr |
Most programs and libraries (/usr/bin, /usr/lib, /usr/local) |
Read-only recommended | Watch for changes |
/etc |
System configuration files (text) | Yes | /etc/passwd, /etc/shadow, /etc/cron* — top targets |
/var |
Variable data: /var/log, /var/spool, /var/lib |
Yes | Logs live here; integrity must be protected |
/tmp |
Temporary, cleared on reboot, usually sticky bit |
Yes (anyone) | Target of race-condition / symlink attacks |
/var/tmp |
Temporary but preserved across reboots | Yes | Persistent payloads |
/home |
User directories | Yes (the owner) | .ssh/, .bash_history |
/root |
root's home | root | — |
/proc |
Pseudo-FS: kernel/process state | Partly | /proc/<pid>/maps, /environ leak secrets |
/sys |
Pseudo-FS: device/kernel objects (sysfs) | Partly | cgroups, modules |
/dev |
Device nodes (devtmpfs) | — | /dev/mem, /dev/kmem are highly sensitive |
/boot |
Kernel, initramfs, bootloader | Rarely | Tampering = rootkit |
/run |
Runtime state (tmpfs, lost on reboot) | Yes | PID files, sockets |
/opt |
Third-party software | Yes | — |
/mnt, /media |
Mount points | — | — |
stat -f / # show the filesystem containing /
findmnt -t ext4,xfs # list mounts with options (ro, nosuid, nodev)
Security note: mounting /tmp, /var/tmp, and /home with the nosuid,nodev,noexec flags is a basic hardening measure — it disables SUID, device nodes, and binary execution from user-writable areas.
2.3. The permission model: 12 bits, reading ls -l character by character, octal, umask
2.3.1. The 12-bit permission structure
Every inode stores a 16-bit st_mode field; the low 12 bits are the access permissions (the remaining 4 high bits encode the file type). The 12 bits are divided into four 3-bit groups:
bit: 11 10 9 | 8 7 6 | 5 4 3 | 2 1 0
SUID SGID STK| r w x | r w x | r w x
<--special-->|<owner>|<group>|<other>
| Group | Bit | Name | Octal value | Meaning on a FILE | Meaning on a DIRECTORY |
|---|---|---|---|---|---|
| Special | 11 | SUID | 4000 | Run with the file owner's UID | (meaningless) |
| Special | 10 | SGID | 2000 | Run with the owning group's GID | New files inherit the directory's GID |
| Special | 9 | Sticky | 1000 | (historically: keep text in swap; now obsolete) | Only the file owner may delete a file in the dir |
| Owner | 8/7/6 | r/w/x | 0400/0200/0100 | read / write / execute | list / create-delete / cd into |
| Group | 5/4/3 | r/w/x | 0040/0020/0010 | same | same |
| Other | 2/1/0 | r/w/x | 0004/0002/0001 | same | same |
Why is x on a directory different from r? r allows reading the list of file names; x allows "traversing" (passing through) to access a file already known by name inside it. You can have x without r: you can reach /dir/file if you know its name, but you cannot ls the directory.
2.3.2. Reading ls -l character by character
-rwxr-xr-- 1 root staff 8192 Jun 19 10:00 tool
drwxr-x--- 2 alice alice 4096 Jun 19 10:00 secret
crw-rw---- 1 root tty 5, 0 Jun 19 10:00 /dev/tty
-rwsr-xr-x 1 root root 55672 Jun 19 10:00 /usr/bin/passwd
The first 10-character string, for example -rwsr-xr-x:
| Position | Character | Meaning |
|---|---|---|
| 1 | - |
File type: -=regular file, d=dir, l=symlink, c=char dev, b=block dev, s=socket, p=named pipe (FIFO) |
| 2–4 | rws |
Owner: r, w, and s = SUID set + x set (if SUID is set but x is off, it shows a capital S) |
| 5–7 | r-x |
Group |
| 8–10 | r-x |
Other |
Special-character rules:
- Owner execute column: x+SUID → s; SUID without x → S.
- Group execute column: x+SGID → s; SGID without x → S.
- Other execute column: x+sticky → t; sticky without x → T.
Example, /tmp:
drwxrwxrwt 18 root root 4096 Jun 19 11:00 /tmp
^ the letter 't' = sticky bit
The sticky bit on /tmp prevents user A from deleting user B's files even though /tmp has w for everyone.
With a device node crw-rw----, the "size" column shows the major, minor numbers (e.g., 5, 0) instead of a byte size — because c/b are devices.
2.3.3. Octal and chmod
chmod 4755 /usr/local/bin/myprog # SUID + rwxr-xr-x
chmod u+s,g-w file # symbolic notation
chmod 1777 /tmp # sticky + rwx for everyone
stat -c '%a %A %U:%G' /usr/bin/passwd
# 4755 -rwsr-xr-x root:root
%a prints octal, %A prints the rwx string — handy for bulk audits.
Hunting SUID/SGID — a mandatory investigative technique:
find / -xdev \( -perm -4000 -o -perm -2000 \) -type f -printf '%M %u %p\n' 2>/dev/null
-xdev: do not cross into other filesystems (avoids scanning NFS, /proc).-perm -4000: matches when at least the SUID bit is set (the-sign means "contains these bits").-printf '%M %u %p\n': prints the mode, owner, path.
Sample output:
-rwsr-xr-x root /usr/bin/sudo
-rwsr-xr-x root /usr/bin/passwd
-rwsr-xr-x root /usr/bin/su
Security note: every SUID-root binary is a privilege-escalation surface. A SUID binary that spawns a shell, reads arbitrary files, or writes arbitrary files can be abused (see the GTFOBins project). Establish a baseline list of SUID binaries and alert when a new entry appears.
2.3.4. umask
umask is a removal mask of permissions for newly created files/directories. The final permission = the default requested permission AND NOT(umask).
- Default for a file:
0666(a new file never gets the x bit automatically). - Default for a directory:
0777.
| umask | New file | New dir | Interpretation |
|---|---|---|---|
022 |
644 |
755 |
other/group cannot write |
027 |
640 |
750 |
other has no access at all |
077 |
600 |
700 |
owner only |
umask # print the current value, e.g. 0022
umask 027
touch a; mkdir b; stat -c '%a %n' a b
# 640 a
# 750 b
Note: umask is inherited from the shell/PAM (/etc/login.defs UMASK, pam_umask). A service running with a loose umask may create world-readable log/secret files. Server hardening typically sets 027 or 077.
2.3.5. ACLs — POSIX Access Control Lists
The 3-class rwx model is insufficient when you need "user X has separate permissions beyond owner/group". ACLs add fine-grained entries, stored in the extended attribute system.posix_acl_access.
setfacl -m u:bob:rwx,g:devs:r-x file.txt # grant bob rwx, the devs group r-x
setfacl -d -m u:bob:rwx /shared # default ACL: new files in the dir inherit it
getfacl file.txt
getfacl output:
# file: file.txt
# owner: alice
# group: alice
user::rw-
user:bob:rwx <- explicit ACL entry
group::r--
mask::rwx <- mask: the ceiling of permissions for named users/groups
other::r--
The mask is important: the effective permissions of user:bob = the ACL entry AND the mask. If setfacl -m m::r--, bob is limited to r-- even though the entry says rwx. When an ACL is present, ls -l shows a + sign:
-rw-rwxr--+ 1 alice alice 0 Jun 19 file.txt
The "group" column in ls -l now shows the mask, not the actual group permissions — a common source of confusion during audits.
Security note: ACLs do not appear in a basic ls -l (only the + sign). Scanning permissions with find -perm alone will miss grants made through ACLs. Use getfacl -R when investigating sensitive permissions.
2.4. /etc/passwd, /etc/shadow, password hashes
2.4.1. /etc/passwd — 7 fields, separated by :
root:x:0:0:root:/root:/bin/bash
sshd:x:106:65534::/run/sshd:/usr/sbin/nologin
alice:x:1000:1000:Alice Nguyen,,,:/home/alice:/bin/bash
| # | Field | Example | Meaning |
|---|---|---|---|
| 1 | username | alice |
Login name |
| 2 | password | x |
x = hash is in /etc/shadow; */! = locked; empty field = no password required (DANGEROUS) |
| 3 | UID | 1000 |
User ID. 0 = root; 1–999 system; ≥1000 ordinary users (depends on login.defs) |
| 4 | GID | 1000 |
Primary group ID |
| 5 | GECOS | Alice Nguyen,,, |
Full name/comment (sub-fields separated by commas) |
| 6 | home | /home/alice |
Home directory |
| 7 | shell | /bin/bash |
Login shell; /usr/sbin/nologin or /bin/false to block login |
Why split out shadow? /etc/passwd must be world-readable (644) so any process can map UID↔name; if hashes were here, anyone could read them for offline cracking. Hashes are moved to /etc/shadow, which only root can read.
getent passwd alice # query through NSS (includes LDAP/SSSD), not just the file
awk -F: '$3==0 {print $1}' /etc/passwd # find every UID-0 account (only root should exist)
Security note: multiple UID-0 accounts = multiple "hidden roots". An empty password field (field 2 blank) allows passwordless login. Both are classic indicators of compromise.
2.4.2. /etc/shadow — 9 fields
Typical permissions 640 root:shadow (or 600).
alice:$6$xQk2...salt...$hashpart...:19800:0:99999:7:14:20000:
| # | Field | Example | Meaning |
|---|---|---|---|
| 1 | username | alice |
Matches passwd |
| 2 | password hash | $6$salt$hash |
Hash or a special state (see below) |
| 3 | last change | 19800 |
Date the password was last changed, in days since 1970-01-01 (epoch days) |
| 4 | min | 0 |
Minimum number of days before it may be changed again |
| 5 | max | 99999 |
Maximum number of days the password remains valid |
| 6 | warn | 7 |
Warning before expiry (days) |
| 7 | inactive | 14 |
Number of days after expiry during which login is still allowed |
| 8 | expire | 20000 |
Date the account is fully disabled (epoch days) |
| 9 | reserved | (empty) | Reserved |
Special states of field 2:
- * or ! → the account cannot log in with a password.
- !$6$... → the password is locked (passwd -l prepends ! to the hash); removing the ! unlocks it.
- Empty → login requires no password.
2.4.3. The $id$salt$hash hash format
The hash field uses the Modular Crypt Format (MCF) syntax: $id$[params]$salt$hash.
id |
Algorithm | Notes |
|---|---|---|
1 |
MD5-crypt | Weak, obsolete |
2a/2b/2y |
bcrypt | Strong; common in apps, rare in shadow |
5 |
SHA-256 crypt | Optional rounds= |
6 |
SHA-512 crypt | Default on many Linux distributions |
y |
yescrypt | Default on Debian 11+/recent Ubuntu; memory-hard, the strongest of the group |
Example of dissecting a SHA-512 record:
$6$rounds=656000$YxZ.Hk1aB2c3D4e$M9...very.long...hashbase64...
| | | |
| | | +-- hash (crypt's special base64 encoding)
| | +------------------- salt (up to 16 characters)
| +--------------------------------- optional parameter (rounds)
+------------------------------------ algorithm id = 6 (SHA-512)
Why a salt? A random salt makes two users with the same password have different hashes, defeating rainbow tables. Why rounds? Increasing the computational cost makes offline brute-force/guessing slower.
# Generate a test hash (yescrypt if the system supports it)
openssl passwd -6 'MatKhau!' # SHA-512
mkpasswd -m yescrypt 'MatKhau!' # mkpasswd is provided by the whois package
chage -l alice # view the password-aging policy (reads shadow)
Security note: MD5 hashes ($1$) must be upgraded. If /etc/shadow is leaked, an attacker runs hashcat/john offline; hashcat modes: 1800 for $6$, 1700 for raw SHA-512, and yescrypt requires a recent hashcat version. Defense: a strong password policy + detection of unauthorized reads of /etc/shadow via auditd.
2.4.4. sudo, /etc/sudoers
sudo lets you run a command with another user's privileges (root by default) without sharing the root password, while logging every command — that is its reason for existing compared with su.
The line syntax in /etc/sudoers (always edit with visudo to check syntax before saving):
user host = (runas_user:runas_group) [TAG:] command
Examples:
# who where (as-whom) what
root ALL=(ALL:ALL) ALL
%admin ALL=(ALL) ALL
alice web01=(www-data) /usr/bin/systemctl restart nginx
bob ALL=(root) NOPASSWD: /usr/bin/journalctl
%dev ALL=(ALL) /usr/bin/apt update, /usr/bin/apt upgrade
Breaking down alice web01=(www-data) /usr/bin/systemctl restart nginx:
- alice: the subject.
- web01: applies only on the host named web01.
- (www-data): runs as the user www-data.
- the permitted command: exactly systemctl restart nginx.
Common TAGs: NOPASSWD: (do not prompt for a password), NOEXEC: (prevent the binary from spawning child commands).
sudo -l # list the current user's sudo permissions
sudo -ll # more detail
Security note: overly broad rules are easily abused. (ALL) NOPASSWD: /usr/bin/vim allows :!sh to become root. Avoid wildcards and editors/interpreters in sudoers. Place drop-in files in /etc/sudoers.d/ (must be 440, with no whitespace in the name).
2.4.5. PAM — Pluggable Authentication Modules
PAM separates authentication logic from the application: sshd, login, and sudo do not code their own password checks but call PAM, configured in /etc/pam.d/<service>. Each line:
<type> <control> <module> [arguments]
type |
Role |
|---|---|
auth |
Verify identity (password, token) |
account |
Check the account is valid (expiry, login hours) |
password |
Change a credential (update the password) |
session |
Set up/tear down a session (mount home, logging, ulimit) |
control |
Behavior |
|---|---|
requisite |
Fail → stop immediately and return failure |
required |
Fail → record the failure but still run the rest of the stack (does not reveal which module failed) |
sufficient |
Success → return success immediately (if no prior required has failed) |
optional |
The result is usually ignored |
[success=1 default=ignore] |
Advanced step-skipping control syntax |
Example /etc/pam.d/sshd (abbreviated) with lockout hardening:
auth required pam_faillock.so preauth silent deny=5 unlock_time=900
auth [success=1 default=bad] pam_unix.so
auth [default=die] pam_faillock.so authfail deny=5 unlock_time=900
account required pam_faillock.so
session required pam_limits.so
pam_faillock locks the account after 5 failures within the window and reopens it after 900 seconds. pam_limits applies /etc/security/limits.conf (limits on processes, file descriptors — protects against fork bombs).
Security note: an incorrect sufficient/required order can create an authentication bypass. Add pam_pwquality to enforce password complexity. An unfamiliar module in /etc/pam.d/ (e.g., a PAM backdoor that writes passwords to a file) is a persistence technique worth scrutinizing.
2.5. Processes
2.5.1. Identity and attributes
Each process has a task_struct in the kernel. The core attributes:
| Attribute | Meaning |
|---|---|
| PID | Process ID (1 = init/systemd) |
| PPID | Parent PID |
| RUID/EUID | Real / Effective UID — EUID determines permissions; SUID makes EUID differ from RUID |
| RGID/EGID | Real / Effective GID |
| SUID/SGID (saved) | UID/GID saved so privileges can be temporarily dropped and regained |
| Supplementary groups | List of secondary groups |
ps -eo pid,ppid,ruid,euid,stat,comm
id alice
cat /proc/self/status | grep -E '^(Uid|Gid|Groups):'
2.5.2. Process state (the STAT column)
| Character | Name | Meaning |
|---|---|---|
R |
Running/Runnable | Running on the CPU or ready to run |
S |
Interruptible sleep | Sleeping while awaiting an event; can be woken by a signal |
D |
Uninterruptible sleep | Sleeping inside kernel I/O; does NOT receive signals (even kill -9) |
T |
Stopped | Stopped by SIGSTOP/SIGTSTP or under debugging |
t |
Traced | Stopped by a debugger |
Z |
Zombie | Dead, waiting for the parent to call wait() to reap the exit code |
X |
Dead | (rarely seen) |
Suffixes in ps (the STAT column): s=session leader, +=foreground group, l=multi-threaded, <=high priority, N=low priority (nice).
Why can't D be killed? The process is in the kernel's I/O path; waking it with a signal could corrupt the device/filesystem state. A prolonged D often signals a hung NFS or a failing disk.
Zombies: consume no resources beyond one entry in the process table; many Z entries mean the parent process is not calling wait(). To kill a zombie, kill/fix the parent process.
2.5.3. The signal table
A signal is an asynchronous notification mechanism. Common numbers (common x86/ARM architectures — a few differ on alpha/mips, verify with kill -l):
| Number | Name | Default | Catchable/blockable? | Description |
|---|---|---|---|---|
| 1 | SIGHUP | Terminate | Yes | Terminal lost; by convention "reload config" for daemons |
| 2 | SIGINT | Terminate | Yes | Ctrl-C |
| 3 | SIGQUIT | Core dump | Yes | Ctrl-\ |
| 9 | SIGKILL | Terminate | NO | Forced kill; cannot be caught/blocked/ignored |
| 11 | SIGSEGV | Core dump | Yes | Invalid memory access |
| 13 | SIGPIPE | Terminate | Yes | Write to a pipe with no reader |
| 15 | SIGTERM | Terminate | Yes | Polite termination request (the default of kill) |
| 17 | SIGCHLD | Ignore | Yes | A child changed state |
| 18 | SIGCONT | Continue | — | Resume a stopped process |
| 19 | SIGSTOP | Stop | NO | Stop the process; cannot be caught |
| 20 | SIGTSTP | Stop | Yes | Ctrl-Z |
kill -l # list all signal names/numbers
kill -TERM 1234 # send SIGTERM
kill -HUP $(pidof nginx) # reload nginx with no downtime
kill -9 1234 # SIGKILL (only when necessary)
Why can't SIGKILL/SIGSTOP be caught? So the admin/kernel always has a way to stop an unruly process; if they could be caught, malware could protect itself indefinitely.
2.5.4. fork() + execve() — how a process is born
parent
| fork() -> create a copy-on-write copy of the page table; returns the child PID to the parent, 0 to the child
+--> child (the copy)
| execve("/bin/ls", argv, envp)
| -> replace the entire memory image with /bin/ls, KEEPING the PID & open fds
v
/bin/ls running
fork()duplicates the process; thanks to copy-on-write, memory is only truly copied when one side writes → fork is cheap.execve()loads a new program over the current address space; open file descriptors are still inherited unless theO_CLOEXECflag is set. This is the basis of shell redirection (section 2.10).- After the child exits, the parent calls
wait()/waitpid()to retrieve the exit code; until it does, the child becomes a zombie.
Security note: a fork+execve("/bin/sh") chain from a non-shell process (e.g., a web server, a daemon) is a classic RCE indicator — write EDR/auditd rules for an execve of a shell whose parent is a network service.
2.5.5. /proc — a window into the kernel and processes
/proc/<pid>/ is a pseudo-filesystem generated dynamically by the kernel.
| Path | Contents | Investigative value |
|---|---|---|
/proc/<pid>/cmdline |
Full command line (arguments separated by NUL \0) |
The real command even if ps argv is spoofed |
/proc/<pid>/exe |
Symlink to the executable binary | Detects deleted binaries ((deleted)) — fileless malware |
/proc/<pid>/cwd |
Symlink to the working directory | — |
/proc/<pid>/environ |
Environment variables (NUL-separated) | Leaks secrets/tokens |
/proc/<pid>/maps |
Mapped memory regions (address, permissions, backing file) | Detects code injection |
/proc/<pid>/fd/ |
Symlinks to every open file descriptor | Find sockets, the log file being written |
/proc/<pid>/status |
UID/GID, capabilities, seccomp, namespace | Privilege audit |
/proc/<pid>/root |
Symlink to the process's root filesystem (differs under chroot/container) | Detects chroot |
tr '\0' ' ' < /proc/$$/cmdline; echo # read cmdline, turn NUL into spaces
ls -l /proc/$(pidof nginx | cut -d' ' -f1)/exe
grep -E 'Cap(Eff|Prm)|Seccomp' /proc/self/status
Detect a deleted-but-still-running binary (a malware hiding technique):
ls -l /proc/*/exe 2>/dev/null | grep deleted
2.5.6. Namespaces — the foundation of containers
A namespace virtualizes one type of kernel resource so that processes inside the namespace believe they own it privately.
| Namespace | Isolates | clone()/unshare parameter |
|---|---|---|
| PID | The PID tree (processes in the NS see their own PID 1) | CLONE_NEWPID |
| NET | Network interfaces, routing table, iptables, ports | CLONE_NEWNET |
| MNT | The mount table | CLONE_NEWNS |
| UTS | hostname, domainname | CLONE_NEWUTS |
| IPC | SysV IPC, POSIX message queues | CLONE_NEWIPC |
| USER | UID/GID mapping (root in the NS = unprivileged outside) | CLONE_NEWUSER |
| CGROUP | The root of the visible cgroup tree | CLONE_NEWCGROUP |
| TIME | The boottime/monotonic clock | CLONE_NEWTIME |
lsns # list every namespace and its owning process
unshare --net --pid --fork --mount-proc bash # create a shell in new net+pid namespaces
readlink /proc/self/ns/net # print the namespace inode, e.g. net:[4026531992]
nsenter -t <pid> -n ss -tlnp # enter a container's net namespace to view its sockets
Why does this matter for security? A container = namespaces + cgroups + capabilities + seccomp/LSM. CLONE_NEWUSER enables "rootless containers" but has historically been the source of many privilege-escalation CVEs. When investigating a container, use nsenter to inspect it from the host without needing a shell inside the container.
2.5.7. cgroups — limiting and measuring resources
cgroups (v2 is the default on modern systems) group processes into a tree and apply CPU/RAM/IO limits. cgroup v2 mounts at /sys/fs/cgroup as a single unified tree.
systemd-cgls # the cgroup tree by systemd unit
cat /sys/fs/cgroup/system.slice/nginx.service/memory.max
cat /sys/fs/cgroup/.../cpu.max # e.g. "200000 100000" = 2 CPUs (quota/period in microseconds)
memory.max sets a RAM ceiling; exceeding it → an OOM kill within that cgroup. pids.max blocks fork bombs. systemd exposes these limits through MemoryMax=, CPUQuota=, and TasksMax= in the unit.
Security note: cgroups are a resource-exhaustion-prevention mechanism (availability), not security isolation like namespaces. Set TasksMax=/MemoryMax= for public-facing services so an abused service cannot bring down the whole host.
2.6. systemd — init, units, services, timers
systemd is PID 1 on most modern distributions, managing the service lifecycle through units. Why replace sysvinit? Parallel startup according to dependencies (faster), process supervision (auto-restart), socket/timer activation, monitoring via cgroups, and structured logging (journald).
2.6.1. The structure of a service unit
File /etc/systemd/system/myapp.service:
[Unit]
Description=My App API
After=network-online.target postgresql.service
Wants=network-online.target
Requires=postgresql.service
[Service]
Type=notify
User=myapp
Group=myapp
ExecStart=/usr/local/bin/myapp --port 8080
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5s
# Hardening:
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
PrivateTmp=true
ReadWritePaths=/var/lib/myapp
CapabilityBoundingSet=
SystemCallFilter=@system-service
MemoryMax=512M
TasksMax=256
[Install]
WantedBy=multi-user.target
| Section | Directive | Meaning |
|---|---|---|
[Unit] |
After= |
Startup order (does not create a hard dependency) |
Requires= |
Hard dependency: if the dependency fails, this unit fails too | |
Wants= |
Soft dependency (recommended, not required) | |
[Service] |
Type= |
Startup model (see the table below) |
ExecStart= |
The main command to run | |
Restart= |
no/on-failure/always/on-abnormal |
|
User=/Group= |
Drop privileges — do NOT run as root if not needed | |
[Install] |
WantedBy= |
The target that "pulls in" the unit when enabled |
The Type= values:
| Type | When systemd considers it "fully started" |
|---|---|
simple |
Right after ExecStart is forked (the default if no Type) |
exec |
After the binary actually execves successfully |
forking |
When the parent process exits (old-style daemons that self-background) — needs PIDFile= |
oneshot |
The process runs to completion and exits (a setup job); usually paired with RemainAfterExit=yes |
notify |
When the process sends sd_notify(READY=1) over a socket — the most accurate |
dbus |
When the service acquires a name on D-Bus |
systemctl daemon-reload # reload after editing a unit
systemctl enable --now myapp.service # enable at boot + start now
systemctl status myapp.service
systemctl cat myapp.service # print the effective unit (including drop-ins)
systemd-analyze security myapp.service # score the hardening (exposure score)
Security note: systemd-analyze security gives an exposure score from 0–10; directives such as ProtectSystem=strict, PrivateTmp=, NoNewPrivileges=, SystemCallFilter=, and CapabilityBoundingSet= lower the score. This is a highly effective way to harden a service without needing a container.
2.6.2. Timer units (replacing cron)
backup.timer:
[Unit]
Description=Nightly backup
[Timer]
OnCalendar=*-*-* 02:30:00
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.target
Paired with backup.service (Type=oneshot). OnCalendar uses the syntax DOW YYYY-MM-DD HH:MM:SS. Persistent=true runs a missed job if the machine was off when it was due. RandomizedDelaySec spreads out the load.
systemctl list-timers --all # see the next and most recent runs
Security advantages over cron: runs in a cgroup, logs to journald, and inherits all of the [Service] hardening — which cron does not have.
2.7. Logging: rsyslog, journald, auth.log, logrotate
2.7.1. The syslog model: facility + severity
Every syslog message carries a PRI = facility×8 + severity. On the wire (RFC 5424) the PRI sits inside < > at the start of the packet.
Facility (selected):
| Number | Facility |
|---|---|
| 0 | kern |
| 1 | user |
| 2 | |
| 3 | daemon |
| 4 | auth (security/authorization) |
| 5 | syslog |
| 10 | authpriv (sensitive auth) |
| 16–23 | local0–local7 (application-defined) |
Severity (lower = more severe):
| Number | Name | Meaning |
|---|---|---|
| 0 | emerg | System is unusable |
| 1 | alert | Action needed immediately |
| 2 | crit | Critical |
| 3 | err | Error |
| 4 | warning | Warning |
| 5 | notice | Normal but noteworthy |
| 6 | info | Informational |
| 7 | debug | Debugging |
Example PRI calculation: facility authpriv(10) + severity info(6) = 10×8+6 = 86 → the packet begins with <86>.
RFC 5424 format:
<PRI>VERSION TIMESTAMP HOSTNAME APP-NAME PROCID MSGID [STRUCTURED-DATA] MSG
<86>1 2026-06-19T02:30:01.003Z web01 sshd 1234 - - Accepted publickey for alice
| Field | Example | Meaning |
|---|---|---|
| PRI | <86> |
facility×8+severity |
| VERSION | 1 |
Protocol version |
| TIMESTAMP | 2026-06-19T02:30:01.003Z |
ISO 8601, with milliseconds + offset/Z |
| HOSTNAME | web01 |
The emitting machine |
| APP-NAME | sshd |
The application |
| PROCID | 1234 |
Usually the PID |
| MSGID | - |
Message type (- = none) |
| STRUCTURED-DATA | - |
Standard key=value pairs |
| MSG | Accepted publickey... |
The content |
2.7.2. rsyslog — filtering & forwarding configuration
/etc/rsyslog.d/50-default.conf (traditional syntax facility.severity destination):
auth,authpriv.* /var/log/auth.log
*.info;mail.none;authpriv.none /var/log/syslog
*.emerg :omusrmsg:*
# Forward to a SIEM over TCP (RELP/TLS recommended for production)
*.* @@siem.internal:6514
auth,authpriv.*→ all severities of these two facilities.mail.none→ exclude the mail facility.@@host:port= TCP (a single@= UDP, which loses packets under high load).
logger -p authpriv.warning "Test message from logger" # inject one message
systemctl restart rsyslog
Security note: forward logs to a centralized SIEM as early as possible — an attacker deleting local logs cannot delete a copy that has left the machine. Prefer TCP/TLS (RELP) so no events are lost and so the transport is secured.
2.7.3. journald
systemd-journald stores logs in a binary, structured, indexed form. Each entry is a set of fields (including trusted fields supplied by the kernel such as _UID, _PID, _SYSTEMD_UNIT — which applications cannot spoof).
journalctl -u sshd.service # logs of one unit
journalctl -p err -b # severity >= err, since this boot
journalctl --since "2026-06-19 02:00" --until "02:30"
journalctl _UID=1000 -o json-pretty # filter by a trusted field, output JSON
journalctl -k # kernel logs (dmesg)
journalctl -f # follow in real time
Persistent journal: by default some distributions store it in /run/log/journal (volatile, lost on reboot). Set Storage=persistent in /etc/systemd/journald.conf and create /var/log/journal to retain it across reboots — mandatory for incident investigation.
Security note: enable Forward Secure Sealing (FSS) to detect log tampering:
journalctl --setup-keys # create a sealing key; an attacker editing the journal is detected at verify time
journalctl --verify
2.7.4. /var/log/auth.log — reading and correlating
Representative lines (Debian/Ubuntu; on RHEL it is /var/log/secure):
Jun 19 02:30:01 web01 sshd[1234]: Accepted publickey for alice from 203.0.113.5 port 51514 ssh2: ED25519 SHA256:abc...
Jun 19 02:31:10 web01 sshd[1240]: Failed password for invalid user admin from 198.51.100.9 port 40222 ssh2
Jun 19 02:32:00 web01 sudo: alice : TTY=pts/0 ; PWD=/home/alice ; USER=root ; COMMAND=/usr/bin/apt update
Jun 19 02:33:00 web01 sshd[1255]: Disconnected from authenticating user root 198.51.100.9 port 40250 [preauth]
Accepted publickey ... ED25519 SHA256:...→ the fingerprint of the login key (cross-check against an allowlist).Failed password for invalid user admin→ a non-existent name = brute-force scanning.sudo: alice : ... COMMAND=→ audit of a privileged command.
# Top IPs causing failed logins
grep "Failed password" /var/log/auth.log \
| grep -oE 'from [0-9.]+' | awk '{print $2}' | sort | uniq -c | sort -rn | head
2.7.5. logrotate
Prevents logs from filling the disk and archives them with a lifecycle. /etc/logrotate.d/nginx:
/var/log/nginx/*.log {
daily
rotate 14
compress
delaycompress
missingok
notifempty
create 0640 www-data adm
sharedscripts
postrotate
[ -f /run/nginx.pid ] && kill -USR1 $(cat /run/nginx.pid)
endscript
}
| Directive | Meaning |
|---|---|
daily |
Rotate every day (weekly/monthly/size 100M) |
rotate 14 |
Keep 14 old copies, then delete |
compress/delaycompress |
Compress old copies; defer compressing the most recent one (so a still-open fd can keep writing) |
create 0640 www-data adm |
Create the new log with the specified permissions/owner |
postrotate ... kill -USR1 |
Tell nginx to reopen the log file (since nginx still holds the old fd pointing to the renamed file) |
Why is postrotate/USR1 needed? A process holds the inode through a file descriptor; renaming the file does not change the inode it is writing to → new logs would fall into the old, renamed file. The USR1 signal (nginx's convention) forces it to reopen the log. Security note: too small a rotate value, or deleting too quickly, can destroy evidence — calibrate to the retention policy and push to a SIEM before deleting.
2.8. Package management: apt and dnf
2.8.1. apt (Debian/Ubuntu, .deb)
apt update # refresh the package list from the repo (download Release/Packages)
apt full-upgrade # upgrade, allowing package removal if needed to resolve dependencies
apt install --no-install-recommends nginx=1.24.0-1
apt-mark hold nginx # pin the version
apt list --installed
dpkg -l | grep nginx # query the low-level package DB
dpkg -V # verify the checksums of installed files (detect tampering)
apt-get -s upgrade # simulate, do not execute
The trust mechanism: the repo has a Release file signed with GPG (Release.gpg/InRelease); the public key lives in /etc/apt/trusted.gpg.d/ or is referenced via signed-by= in the .sources file. apt checks signatures and hashes before installing → protecting against forged packages over the network.
2.8.2. dnf (RHEL/Fedora/Rocky, .rpm)
dnf check-update
dnf install nginx
dnf history # view the transaction history
dnf history undo <id> # roll back a transaction
dnf needs-restarting -r # services needing a restart after the update (dnf-utils package)
rpm -qa | grep nginx
rpm -V nginx # verify: columns S(size) M(mode) 5(md5) ... differing from the baseline
rpm -qf /usr/sbin/nginx # which package this file belongs to
RPM signs each package with GPG; gpgcheck=1 in the .repo enforces the check.
Security note for both: use only signed HTTPS repos; pin versions (hold/version lock) for critical systems; dpkg -V / rpm -V are quick integrity-checking tools to detect replaced binaries. Set up unattended-upgrades (Debian) / dnf-automatic for automatic security patches.
2.9. cron — 5 time fields
A user's crontab -e, or the system files /etc/crontab and /etc/cron.d/* (which have an additional 6th USER field).
┌──────── minute (0–59)
│ ┌────── hour (0–23)
│ │ ┌──── day of month (1–31)
│ │ │ ┌── month (1–12)
│ │ │ │ ┌ day of week (0–7; both 0 and 7 = Sunday)
│ │ │ │ │
* * * * * command
| Field | Range | Special characters |
|---|---|---|
| minute | 0–59 | * any value; */5 every 5; 1,15 a list; 0-30 a range |
| hour | 0–23 | as above |
| day | 1–31 | as above |
| month | 1–12 | or the names jan–dec |
| DOW | 0–7 | or sun–sat |
Examples:
*/5 * * * * /usr/local/bin/health-check.sh # every 5 minutes
0 2 * * 1-5 /usr/local/bin/backup.sh # 02:00 Mon–Fri
30 3 1 * * root /usr/local/bin/monthly.sh # (in /etc/cron.d) the 6th field = user
@reboot /usr/local/bin/startup.sh # at boot (macro)
Macros: @reboot @daily @hourly @weekly @monthly @yearly.
Security note: cron is a favored persistence point. Inspect: every user's crontab -l (/var/spool/cron/crontabs/* on Debian, /var/spool/cron/* on RHEL), /etc/crontab, and /etc/cron.{d,hourly,daily,weekly,monthly}. Entries with @reboot, unfamiliar download/encode commands, or that point to /tmp are highly suspicious. cron logs to syslog (the cron facility) — correlate run times with suspicious activity.
2.10. Bash: file descriptors, redirection, pipes
2.10.1. The three standard file descriptors
Every process opens three fds; they are just integers pointing into the kernel's fd table:
| fd | Name | Default |
|---|---|---|
| 0 | stdin | Keyboard / terminal |
| 1 | stdout | Terminal |
| 2 | stderr | Terminal |
Why separate stdout and stderr? So that "clean" results (1) can go one way and errors (2) another — without mixing data with error messages during automated processing.
2.10.2. Redirection — the full table
| Syntax | Behavior |
|---|---|
> file |
stdout overwrites the file (shorthand for 1>) |
>> file |
stdout appends |
2> file |
stderr overwrites |
2>> file |
stderr appends |
&> file / >& file |
both stdout+stderr (a bashism) |
> file 2>&1 |
stdout into the file, THEN stderr points to "wherever 1 currently points" |
2>&1 > file |
(WRONG ORDER) stderr points to the terminal, then stdout is redirected to the file |
< file |
stdin reads from the file |
<<EOF ... EOF |
here-document |
<<<"string" |
here-string |
2>/dev/null |
discard errors |
n>&- |
close fd n |
Why does the order of 2>&1 matter? Redirection is processed left→right. 2>&1 means "fd 2 = a copy of fd 1's current destination". If fd 1 has not been changed (still the terminal), stderr also goes to the terminal. You must put > file FIRST so fd 1 points to the file, then 2>&1 copies the correct destination.
make 2>build-errors.log # split build errors into a separate file
./script.sh > out.log 2>&1 # the correct merge: both go to out.log
./script.sh &>out.log # equivalent (bash)
diff <(sort a.txt) <(sort b.txt) # process substitution: each <() is an fd /dev/fd/N
exec 3>/var/log/app.audit # open a custom fd 3
echo "event" >&3 # write through fd 3
2.10.3. Pipe
cmd1 | cmd2: the kernel creates a pipe (a kernel buffer ~64KB); cmd1's stdout (fd 1) connects to cmd2's stdin (fd 0). The two processes run concurrently; cmd1 blocks when the buffer is full and cmd2 blocks when it is empty — this is the backpressure mechanism.
set -o pipefail # the pipeline's exit code = the first command that fails (not just the last)
cmd1 | cmd2 ; echo ${PIPESTATUS[@]} # an array of each pipe segment's exit code
Note: by default the pipeline's exit code is only that of the last command — grep x file | wc -l returns 0 even if grep found nothing. pipefail is very important in security scripts so errors are not "swallowed".
2.11. Text-processing tools: grep, awk, sed, cut, sort, uniq — practical log analysis
This is the daily log-investigation toolkit. Below is a complete pipeline analyzing auth.log, with each segment explained.
2.11.1. grep — filter lines
grep -E "Failed password|Invalid user" /var/log/auth.log
grep -c "Accepted" /var/log/auth.log # count matching lines
grep -oE '([0-9]{1,3}\.){3}[0-9]{1,3}' file # print ONLY the matching part (IPv4)
grep -v "cron" file # invert: drop matching lines
grep -i -A3 -B1 error log # case-insensitive, print 3 lines after, 1 before
-E enables extended regex, -o prints only the matching string (not the whole line), -v inverts, -c counts, -A/-B/-C give context.
2.11.2. awk — process by field
awk splits each line into fields $1..$NF by FS (whitespace by default). Structure: awk 'PATTERN { ACTION }'.
# Count successful SSH logins per user
awk '/Accepted/ {for(i=1;i<=NF;i++) if($i=="for") print $(i+1)}' /var/log/auth.log \
| sort | uniq -c | sort -rn
/Accepted/: a pattern, processes only lines containing "Accepted".- the
forloop finds the wordforand prints the word right after it (which is the username) — robust even if the columns shift. NF= the number of fields,$(i+1)= the next field.
# Total bytes transferred per IP from an nginx access log (field 1 = IP, field 10 = bytes)
awk '{sum[$1]+=$10} END {for(ip in sum) printf "%-16s %d\n", ip, sum[ip]}' access.log \
| sort -k2 -rn | head
sum[$1]+=$10 uses an associative array; the END block runs after the whole file is read.
2.11.3. sed — stream editor
sed -n '100,150p' big.log # print only lines 100–150
sed 's/[0-9]\{1,3\}\(\.[0-9]\{1,3\}\)\{3\}/[IP]/g' log # anonymize IPs
sed -i.bak 's/PermitRootLogin yes/PermitRootLogin no/' /etc/ssh/sshd_config
sed '/^#/d;/^$/d' config # delete comment lines and blank lines
-n disables automatic printing, p prints; s/old/new/g substitutes globally; -i.bak edits in place and saves a .bak backup; d deletes a line.
2.11.4. cut — cut by column/character
cut -d: -f1,7 /etc/passwd # fields 1 and 7, separated by ':'
cut -d' ' -f1-3 access.log # the first 3 fields
cut -c1-15 /var/log/syslog # the first 15 characters (the timestamp column)
-d sets the delimiter, -f selects fields, -c selects by character.
2.11.5. sort & uniq — sorting and counting duplicates
sort -t: -k3 -n /etc/passwd # sort by UID (field 3), numerically
sort -k2 -rn data # field 2, numeric, descending
uniq -c # count consecutive identical lines (MUST sort first)
uniq -d # print only lines that have duplicates
Why sort before uniq? uniq only collapses adjacent identical lines; you must sort to bring them together first.
2.11.6. A practical pipeline — "Top 10 SSH brute-force IPs"
grep "Failed password" /var/log/auth.log \
| grep -oE 'from ([0-9]{1,3}\.){3}[0-9]{1,3}' \
| awk '{print $2}' \
| sort \
| uniq -c \
| sort -rn \
| head -n 10
Explanation of each stage:
1. grep "Failed password" → keep failed-login lines.
2. grep -oE 'from <IPv4>' → extract exactly the from 198.51.100.9 portion.
3. awk '{print $2}' → drop the word from, leaving the IP.
4. sort → bring identical IPs adjacent (preparing for uniq).
5. uniq -c → count how many times each IP occurs (prepends the count to the line).
6. sort -rn → sort descending by the count.
7. head -n 10 → the 10 most active attacking IPs.
Sample output:
412 198.51.100.9
207 203.0.113.77
95 192.0.2.44
This result feeds directly into an allowlist/blocklist or an alert (see fail2ban, 2.12.2).
2.12. Hardening: sshd, fail2ban, netfilter, SELinux/AppArmor
2.12.1. sshd_config — directive by directive
File /etc/ssh/sshd_config; apply with systemctl reload sshd. A sample hardened configuration:
Port 22
AddressFamily inet
# Authentication
PermitRootLogin no
PubkeyAuthentication yes
PasswordAuthentication no
KbdInteractiveAuthentication no
PermitEmptyPasswords no
MaxAuthTries 3
MaxSessions 4
LoginGraceTime 30
AuthenticationMethods publickey
# User restrictions
AllowGroups ssh-users
AllowUsers alice@203.0.113.0/24
# Cryptography (strong algorithms only)
KexAlgorithms curve25519-sha256,curve25519-sha256@libssh.org
Ciphers chacha20-poly1305@openssh.com,aes256-gcm@openssh.com
MACs hmac-sha2-512-etm@openssh.com,hmac-sha2-256-etm@openssh.com
# Reduce the surface
X11Forwarding no
AllowAgentForwarding no
AllowTcpForwarding no
PermitTunnel no
ClientAliveInterval 300
ClientAliveCountMax 2
LogLevel VERBOSE
| Directive | Security effect |
|---|---|
PermitRootLogin no |
Forces login as an ordinary user then sudo → an audit trail; blocks direct brute-forcing of root |
PasswordAuthentication no |
Public keys only → eliminates password brute-forcing |
MaxAuthTries 3 |
Disconnect after 3 failures |
LoginGraceTime 30 |
Close an unauthenticated connection after 30s (prevents holding a slot) |
AuthenticationMethods publickey |
Can enforce multi-factor: publickey,keyboard-interactive |
AllowGroups/AllowUsers |
Allowlist who may SSH (implicitly deny the rest) |
KexAlgorithms/Ciphers/MACs |
Remove weak algorithms; -etm (encrypt-then-MAC) is safer |
LogLevel VERBOSE |
Logs even the fingerprint of the key used to log in (very useful for investigation) |
sshd -t # check the syntax before reloading (AVOID locking yourself out)
sshd -T | grep -i permitroot # print the actual effective configuration
Security note: always sshd -t and keep an open session while reloading, in case a misconfiguration locks out access. Consider not sending a version banner and using Match blocks for group/address-specific configuration.
2.12.2. fail2ban — dynamic brute-force blocking
fail2ban reads logs, uses a regex (failregex) to detect failures, then invokes an action (typically inserting a firewall rule) to ban the IP for a period.
/etc/fail2ban/jail.local:
[DEFAULT]
bantime = 1h
findtime = 10m
maxretry = 5
banaction = nftables-multiport
ignoreip = 127.0.0.1/8 203.0.113.0/24
[sshd]
enabled = true
port = ssh
backend = systemd
maxretry = 3
bantime = 24h
| Parameter | Meaning |
|---|---|
findtime |
The time window for counting failures (10 minutes) |
maxretry |
The number of failures within findtime to be banned (3) |
bantime |
The ban duration (-1 = permanent); incremental banning can be enabled |
backend = systemd |
Read from journald instead of a log file |
ignoreip |
Allowlist that is never banned |
fail2ban-client status sshd # see currently banned IPs and statistics
fail2ban-client set sshd unbanip 198.51.100.9
fail2ban-regex /var/log/auth.log /etc/fail2ban/filter.d/sshd.conf # test failregex
An example failregex (in filter.d/sshd.conf) matching a failure line:
^.*Failed (?:password|publickey) for .* from <HOST>
<HOST> is a fail2ban macro that it replaces with an IP regex and uses to extract the address to ban.
Security note: set ignoreip for the admin range so you do not lock yourself out. fail2ban defends against brute-force but does NOT replace disabling password auth — combine both.
2.12.3. Netfilter: iptables and nftables
Both configure the kernel's netfilter framework. Packets traverse a set of hooks along this path:
PREROUTING FORWARD POSTROUTING
packet in -> [raw->mangle->nat] --> routing? --yes--> [mangle->filter] --> [mangle->nat] --> out
|
| (destined for this machine)
v
INPUT [mangle->filter] --> local process
|
OUTPUT [...] --> POSTROUTING
The four iptables tables — each has one job, don't mix them:
| Table | What it exists for | Chains present | When you touch it |
|---|---|---|---|
filter |
Allow/block packets — the firewall proper | INPUT, FORWARD, OUTPUT | 90% of the time; the default table when -t is omitted |
nat |
Rewrite addresses/ports (SNAT/DNAT/MASQUERADE) | PREROUTING, OUTPUT, POSTROUTING | The machine acts as a gateway/port-forwarder; only NEW packets traverse it |
mangle |
Modify packet headers (TTL, TOS/DSCP, marks) | All 5 chains | Marking packets for QoS/policy routing — rarely needed |
raw |
Processing BEFORE conntrack (NOTRACK) |
PREROUTING, OUTPUT | Excluding high-volume traffic from the conntrack table |
What is a chain? A list of rules attached to a hook; the first rule a packet matches decides its verdict (ACCEPT/DROP/REJECT/jump to a sub-chain) — if nothing matches, the chain's policy applies. Besides the 5 built-in chains you can create your own (iptables -N ssh-guard) and -j ssh-guard into it to group rules by topic — easier to read and to count (each rule has its own counter, visible with iptables -L -v -n).
The anatomy of a rule, worth internalizing: iptables -t <table> -A <chain> <matches...> -j <verdict> — the matches are AND-ed conditions (-p tcp --dport 22, -s 203.0.113.0/24, -i eth0, -m conntrack --ctstate NEW...), and the verdict is the action when all of them match.
iptables (traditional syntax) — a deny-by-default policy for INPUT:
iptables -P INPUT DROP # default policy: block
iptables -P FORWARD DROP
iptables -P OUTPUT ACCEPT
iptables -A INPUT -i lo -j ACCEPT # allow loopback
iptables -A INPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT # allow established connections
iptables -A INPUT -p tcp --dport 22 -m conntrack --ctstate NEW -m limit --limit 5/min -j ACCEPT
iptables -A INPUT -p tcp --dport 443 -j ACCEPT
iptables -A INPUT -j LOG --log-prefix "DROP_IN: " --log-level 4
Breaking down the SSH rule:
- -A INPUT: append to the INPUT chain.
- -p tcp --dport 22: TCP destined for port 22.
- -m conntrack --ctstate NEW: only packets initiating a new connection.
- -m limit --limit 5/min: at most 5 new connections/minute (anti brute-force/flood).
- -j ACCEPT: allow it through.
Why is the ESTABLISHED,RELATED rule needed? Stateful: once a valid inbound/outbound connection is established, the reply packets in the ESTABLISHED state are allowed through without opening a separate port — the foundation of a stateful firewall.
A few practical rules you end up needing:
# Block one IP / one IP range (insert at the TOP of the chain with -I so it matches before any ACCEPT)
iptables -I INPUT -s 198.51.100.9 -j DROP
iptables -I INPUT -s 203.0.113.0/24 -j DROP
# Sliding-window SSH rate limit using the recent module:
# more than 4 NEW connections to port 22 within 60s from the same source -> drop
iptables -A INPUT -p tcp --dport 22 -m conntrack --ctstate NEW \
-m recent --name ssh --set
iptables -A INPUT -p tcp --dport 22 -m conntrack --ctstate NEW \
-m recent --name ssh --update --seconds 60 --hitcount 4 -j DROP
# View rules with counters (packets/bytes matched) — the quickest way to tell whether a rule is "firing"
iptables -L INPUT -v -n --line-numbers
# Delete a rule by line number
iptables -D INPUT 3
For long IP lists (hundreds/thousands of entries), do not stack sequential rules — every packet has to walk the whole list. Use ipset (with iptables) or an nftables set (example below): O(1) hash lookup, a single rule referencing the whole collection.
The iptables ↔ nftables relationship (read carefully or you'll think you're using one while actually using the other): on modern distributions (Debian 10+, Ubuntu 20.04+, RHEL 8+), the iptables command is usually just a shim that translates the old syntax onto the nftables backend — known as iptables-nft. Check with iptables -V: (nf_tables) means the nft shim, (legacy) means the old backend. The practical consequence: rules created via iptables and rules created via nft live in the same kernel but do not fully show up in each other's tooling — a firewall audit must use nft list ruleset (which sees both), so never run just iptables -L and conclude "this machine has no rules". Docker/Kubernetes and fail2ban still commonly insert rules through the iptables layer — one more reason to audit with nft list ruleset.
nftables (the modern replacement, a single nft tool for IPv4/IPv6) — /etc/nftables.conf:
#!/usr/sbin/nft -f
flush ruleset
table inet filter {
set blocklist {
type ipv4_addr
flags interval
elements = { 198.51.100.9, 203.0.113.0/24 }
}
chain input {
type filter hook input priority 0; policy drop;
iif "lo" accept
ip saddr @blocklist drop
ct state established,related accept
ct state invalid drop
tcp dport 22 ct state new limit rate 5/minute accept
tcp dport 443 accept
ip protocol icmp icmp type echo-request limit rate 1/second accept
log prefix "nft-drop: " counter
}
chain forward { type filter hook forward priority 0; policy drop; }
chain output { type filter hook output priority 0; policy accept; }
}
nft -f /etc/nftables.conf # load (atomic: all or nothing)
nft list ruleset # view all rules (with counters)
nft list table inet filter
nft add element inet filter blocklist { 203.0.113.99 } # add an IP to the set at runtime
Why does nftables replace iptables? A single framework for v4/v6/arp/bridge (an inet table covers IPv4+IPv6 at once — iptables needs a parallel ip6tables), concise set/map syntax (a blocklist of thousands of IPs in one hash-lookup set instead of a sequential walk), atomic reloads (no moment of "half old ruleset, half new"), and easier scripting.
Making rules survive a reboot — iptables/nft rules live only in kernel memory; a reboot wipes them clean:
# The nftables way (recommended): rules are a config FILE, loaded by a service at boot
nft list ruleset > /etc/nftables.conf # or write the file by hand and validate with nft -f
systemctl enable --now nftables
# The traditional iptables way (Debian/Ubuntu)
apt install iptables-persistent # stores at /etc/iptables/rules.v4 and rules.v6
netfilter-persistent save
# RHEL-family: dnf install iptables-services; service iptables save
The nice part of the nftables way: /etc/nftables.conf is a plain text file — put it in git, review diffs like code, run nft -c -f (check) before applying.
Security note: default to policy drop for INPUT/FORWARD; open only the necessary ports; rate-limit the admin port; log dropped packets for investigation. When editing a firewall remotely, protect yourself: work inside tmux/keep a spare SSH session open, or schedule a rollback job like echo 'nft -f /etc/nftables.conf.known-good' | at now + 5 minutes and cancel it once you've confirmed you can still get in — locking yourself out of a server with one DROP rule is a classic mistake.
2.12.4. SELinux and AppArmor — Mandatory Access Control (MAC)
Traditional rwx permissions are DAC (Discretionary): the owner decides, and root bypasses everything. MAC applies a system policy that even root is bound by — if a service is compromised, MAC limits the damage. Both install hooks through LSM in the kernel.
SELinux (RHEL/Fedora) — every process and file has a security context of the form user:role:type:level:
system_u:system_r:httpd_t:s0 <- the httpd process
unconfined_u:object_r:httpd_sys_content_t:s0 <- a web file nginx serves
Decisions are based on type enforcement: the httpd_t domain may only do what the policy permits on the related types.
getenforce # Enforcing / Permissive / Disabled
setenforce 0 # temporarily switch to Permissive (logs only, does not block)
ps -eZ | grep nginx # view a process's context (the -Z flag)
ls -Z /var/www/html # view a file's context
ausearch -m AVC -ts recent # find events blocked by SELinux (AVC denials)
# Fix the context for a custom web directory:
semanage fcontext -a -t httpd_sys_content_t "/srv/www(/.*)?"
restorecon -Rv /srv/www
setsebool -P httpd_can_network_connect on # turn on a boolean (allow httpd to connect out)
An AVC denial in the log:
type=AVC msg=audit(...): avc: denied { read } for pid=1234 comm="nginx"
name="secret.txt" scontext=system_u:system_r:httpd_t:s0
tcontext=unconfined_u:object_r:user_home_t:s0 tclass=file
Read it: the httpd_t domain was denied read on a file of type user_home_t — exactly in the spirit of "a web server must not read home files". Do NOT disable SELinux just "to make it run"; instead fix the context or the boolean.
AppArmor (Ubuntu/SUSE) — profiles by path (easier to read than SELinux). /etc/apparmor.d/usr.sbin.myapp:
#include <tunables/global>
/usr/sbin/myapp {
#include <abstractions/base>
capability net_bind_service,
network inet stream,
/etc/myapp/** r,
/var/lib/myapp/** rw,
/var/log/myapp.log w,
/usr/sbin/myapp mr,
deny /home/** rwx,
}
Each line = a path rule + permissions (r w m=mmap exec ix/px...). deny /home/** rwx absolutely blocks access to home.
aa-status # which profiles are enforcing/complaining
aa-complain /usr/sbin/myapp # log-only mode (learn behavior)
aa-enforce /usr/sbin/myapp # turn on blocking
journalctl | grep apparmor # view denials (ALLOWED/DENIED)
Why path-based vs type-based? AppArmor is easy to write/read (by path) but weak when files are renamed/hardlinked; SELinux attaches labels to the inode, making it stricter but with a steeper learning curve. Both are key defensive layers: when an RCE occurs in a confined service, the payload is blocked at any operations outside the profile/policy.
2.13. Performance & resource diagnostics — the dashboard says "where it's high", the commands say "which process is causing it"
Why is this section in a security notebook? First, availability is one pillar of the CIA triad — a machine that dies from resource exhaustion is an information-security incident too. Second, a lot of intrusion indicators first show up as resource anomalies: a cryptominer = abnormally high CPU, a DoS = load/network spiking, and a full disk means logging stops — the attacker gets a free blind spot. The skill of "spot the abnormal number → find the process behind it" serves operations and investigation alike.
The way this section is organized follows the USE method (Utilization – Saturation – Errors) proposed by Brendan Gregg: walk each resource in turn (CPU, memory, disk, I/O, network) and ask three questions of it — how much is in use, is it saturated, are there errors. On top of that sit two tiers of tooling, each answering a different question. - The metric/dashboard layer (Prometheus + node_exporter, graphed in Grafana — or any equivalent stack): answers "where is it high, since when, and what's the trend". Metrics are aggregates — they don't know which process is responsible. - The on-box command layer (SSH in): answers "which process, which file, which connection is causing it". Commands see individual processes but have no history.
The two layers complement each other: a dashboard alone knows the disease but not the culprit; commands alone capture the current state but not "high since when, is it periodic". The workflow on an alert: the dashboard narrows it to an axis (CPU/RAM/disk/I-O/network/load) → SSH in → run exactly that axis's command group.
The commands below need the
sysstatpackage (providesmpstat/iostat/pidstat/sar), plushtop,iotop,ncdu:apt install sysstat htop iotop ncdu.
2.13.1. The 30-second overview — run before you think
Six commands, in order, for the big picture before digging in:
uptime # load average over 1/5/15 minutes — compare to core count
free -h # RAM: look at the available column, not used
df -h # which mount is filling up
top # (or htop) overall CPU/RAM + the loudest processes
dmesg -T | tail -30 # what the kernel just complained about: OOM kills, I/O errors, segfaults
systemctl --failed # which services are down
These 30 seconds usually already isolate the problem axis; the rest is a deep dive per axis below.
2.13.2. The CPU axis — read by mode, not just by %
"CPU at 90%" by itself says nothing; the right question is which mode that 90% belongs to. mpstat (sysstat package) breaks CPU down by mode, per core:
mpstat -P ALL 1 3 # one sample per second, 3 times, all cores
CPU %usr %nice %sys %iowait %irq %soft %steal %idle
all 72.31 0.00 5.12 0.75 0.00 0.51 9.87 11.44
0 88.00 0.00 4.00 0.00 0.00 1.00 17.00 7.00
| Column | CPU time spent on | Sustained high means |
|---|---|---|
%usr |
Application code (user space) | An app is genuinely burning CPU → find it with htop/pidstat; if you don't recognize the process, suspect a cryptominer |
%sys |
Kernel (syscalls, drivers) | The app is hammering syscalls (tiny I/O, constant forking) — inspect with strace/perf |
%iowait |
Sitting idle waiting for disk I/O | The bottleneck is the DISK, not the CPU — jump to the Disk I/O axis (2.13.5); more CPU won't help |
%steal |
Taken by the hypervisor to serve other VMs | On a cloud VM: noisy neighbor / the provider throttling CPU (burstable credits exhausted). Not fixable from inside the VM — change shape/host or accept it |
%irq/%soft |
Hard/soft interrupts (usually network) | Abnormally high %soft plus heavy traffic → suspect a flood |
%idle |
Idle | — |
%iowait and %steal are the two classic traps: both make "CPU look high" on a dashboard while the culprit isn't the CPU at all. Finding the process by CPU:
ps aux --sort=-%cpu | head -12 # snapshot of right now
pidstat 1 5 # over time — catches processes that "pulse"
pidstat beats ps by sampling continuously: a process that only spikes in bursts shows up instead of slipping between two ps runs.
2.13.3. The RAM axis — PSI instead of "used is high"
The biggest trap on the RAM axis: "used is high" is usually NOT a problem. Linux deliberately uses free RAM as page cache (re-reads of files skip the disk), and many apps (JVM, Elasticsearch, databases) allocate a large heap and simply keep it — that's designed behavior, not a leak. The number to watch in free -h is available (an estimate of RAM that can be handed to new processes immediately, reclaimable cache included), not used.
So when is RAM genuinely short? The kernel answers directly via PSI — Pressure Stall Information (kernel ≥ 4.20): instead of measuring "how much is in use", PSI measures the total time processes are stalled waiting for the resource — i.e., it measures exactly what we care about: "is anyone actually suffering from lack of RAM?"
cat /proc/pressure/memory
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0
some= % of time at least one task was stalled on memory;full= % of time all non-idle tasks were stalled (the whole machine freezes up).avg10/60/300= rolling averages over 10 seconds / 1 minute / 5 minutes.- Reading it: all zeros → RAM is fine no matter what
usedsays.some avg10positive and sustained → shortage beginning;fullpositive → the machine is visibly stalling, typically just before an OOM kill. The same format exists for/proc/pressure/cpuand/proc/pressure/io— node_exporter collects all three, and they make far better alerting panels than % used.
The full check sequence:
free -h # look at available; how much swap is in use
cat /proc/pressure/memory # PSI — the measure of "genuinely short"
vmstat 1 5 # si/so columns: swap-in/out (KB/s) happening right now
ps aux --sort=-%mem | head -12 # top RAM consumers (RSS column)
dmesg -T | grep -iE "oom|killed process" # has the kernel already had to execute someone
si/sopersistently positive = the machine is swapping back and forth (thrashing) — performance falls off a cliff even though RAM "isn't full".- An OOM kill in
dmesgis evidence you're already too late: the kernel kills the process with the highestoom_score. If the victim is an important service, considerMemoryMax=(cgroups, section 2.5.7) on the other services rather than reflexively adding RAM.
2.13.4. The disk-space axis — df lies in two places
df -h # which mount is full — but not enough, keep reading
df -i # INODES: running out also reports "No space left" with tens of GB free
sudo du -xh --max-depth=1 / 2>/dev/null | sort -h | tail -15 # biggest directories (-x: don't wander onto other mounts)
ncdu -x / # like du but interactive — very fast drill-down
sudo lsof +L1 | grep -i deleted # files DELETED but still held open by a process
docker system df # how much Docker is using (images/containers/volumes/build cache)
journalctl --disk-usage # how much the journal is using
The two places df -h "lies":
1. Out of inodes: a filesystem has a finite inode count; millions of small files (session files, caches, mail queues) exhaust the inodes while plenty of space remains — the error is still "No space left on device". df -i exposes it instantly (IUse% at 100%).
2. Deleted files not yet released: deleting a file only removes its name from the directory; the inode and its data persist as long as any process holds an open file descriptor (the same mechanism we met with logrotate, section 2.7.5). The classic symptom: you rm a 20GB log file and df doesn't move. lsof +L1 (lists files with link count 0) names the process; the correct fix is to reload/restart that process (or send a reopen-logs signal such as USR1), not to keep deleting things.
On Docker machines the usual suspect is /var/lib/docker: old images stacking up with every deploy, build cache, container logs. docker system df -v gives the detailed listing. Clean deliberately (docker image prune -a --filter "until=168h" — images only, with a time cutoff) and never swing docker system prune around on a machine with data volumes — one mistaken --volumes flag and real data is gone.
The security angle: full disk = logging stops = investigative blindness; "disk almost full" therefore deserves to be treated as a security alert, not just an ops one. In the other direction, attackers know this too — one form of anti-forensics is deliberately filling the disk before acting.
2.13.5. The disk I/O axis — a "busy" disk is not a "full" disk
A disk with free space can still choke because its I/O bandwidth is saturated — the symptom surfaces as %iowait (2.13.2) and apps that are "slow for no reason".
iostat -xz 1 5 # -x: extended stats; -z: hide silent devices
Device r/s w/s rkB/s wkB/s await %util
sda 1.2 310.5 48.0 42817.3 38.20 97.40
%util= % of time the device had I/O in flight. Near 100% sustained = the disk's schedule is packed. (For SSDs/NVMe, which process in parallel, 100%%utilisn't necessarily the true ceiling — read it together withawait.)await= average time for a request to complete, queueing time included (ms). This is the number the app actually "feels": high%utilwith lowawaitmeans the disk is busy but keeping up; both high means genuine saturation.r/s,w/s,rkB/s,wkB/s: the shape of the load — thousands of small requests or a few big ones, reads or writes.
Finding the culprit:
sudo iotop -o # -o: only show processes CURRENTLY doing I/O
pidstat -d 1 5 # kB read/written per process over time
cat /proc/pressure/io # PSI for I/O — read exactly like memory PSI
Usual suspects: backup/compression jobs running at peak hours, databases scanning whole tables for lack of an index, over-verbose logging, and swap thrashing (in which case the root disease is RAM — go back to 2.13.3).
2.13.6. The network axis — look at drops/errors, not % bandwidth
ss -tulpn # LISTENING ports + processes — inspect the exposed surface
ss -s # connection totals by state (estab, timewait...)
ss -tan state established | wc -l # count open connections
ip -s link # RX/TX counters: errors, dropped, overrun per interface
sar -n DEV 1 5 # per-interface throughput over time
ip -s link prints RX/TX blocks; the columns worth watching are errors/dropped:
RX: bytes packets errors dropped overrun mcast
9.1G 8123456 0 12043 0 0
- Hitting the bandwidth ceiling is rarely the real problem on an ordinary server; steadily climbing drops/errors are the disease (buffer overruns, NIC/driver faults, the host not keeping up → dropped packets, retransmits, rising latency). Alerts belong on the growth rate of the drop counters, not on % bandwidth.
ss -swithestabspiking abnormally from scattered sources is the shape of a DoS/flood; a strange new listening port inss -tulpnis a backdoor indicator (compare against a baseline).
2.13.7. The load axis — high load does not mean short of CPU
Load average on Linux counts both R processes (waiting for CPU) and D processes (uninterruptible — stuck in I/O, section 2.5.2). This is the most commonly misread point: a load of 20 on a 4-core box could be 20 processes fighting for CPU — or an idle CPU and 20 processes hanging on a dead NFS mount.
uptime && nproc # load RELATIVE TO core count — load/core > 1 sustained is what matters
vmstat 1 5 # r column = waiting for CPU; b = blocked; wa = %iowait
ps -eo state,pid,comm | grep "^D" # roll call of D-state processes
Reading the verdict: high load + high r + high %usr → genuinely short of CPU. High load + high b + many D processes + high wa → stuck on I/O (failing disk, hung NFS, throttled cloud storage) — go back to the Disk I/O axis, and remember D processes cannot be killed with kill -9 (2.5.2).
2.13.8. Quick lookup: symptom → first command
| Dashboard/panel alert | Run first | Question being answered |
|---|---|---|
| High CPU | mpstat -P ALL 1, htop |
Which mode? Which process? |
| High RAM / memory PSI | cat /proc/pressure/memory, ps aux --sort=-%mem |
Genuine shortage or just high used? Who's eating it? |
| Disk full | df -h; df -i, du -xh --max-depth=1 /, lsof +L1 |
Out of space, out of inodes, or ghost files? |
| Disk I/O / io PSI | iostat -xz 1, iotop -o |
Saturated yet? Which process is writing? |
| Network | ip -s link, ss -tulpn |
Drops/errors or just lots of traffic? Strange ports? |
| High load / procs blocked | vmstat 1, ps -eo state,pid,comm \| grep "^D" |
Short of CPU or stuck on I/O? |
My own experience: turn this table into an actual internal runbook (which panel is red → paste which block of commands) — at 2 a.m. during an incident nobody remembers iostat's flags.
2.14. Summary of the defensive mindset on Linux
- Privilege boundaries are the root: user/kernel (syscall, seccomp), DAC (rwx/ACL/SUID), MAC (SELinux/AppArmor), capabilities, namespaces. Each layer narrows the damage when the layer above is breached.
- Least privilege: services run as their own user,
NoNewPrivileges, a deny-by-default firewall, narrow sudoers,nosuid/noexecmounts. - Observability: structured logs (journald + auditd), pushed centrally to a SIEM before an attacker deletes them, and fluent reading of
auth.log/AVC with grep/awk. - Integrity:
dpkg -V/rpm -V, FSS for the journal, baselines of SUID/cron to detect changes. - Availability is security too: resource exhaustion (CPU/RAM/disk/I-O) is both an incident and an attack symptom; the two-layer diagnosis "the dashboard says where it's high, the commands say which process" (2.13) is a skill shared by operations and investigation.
- Every hardening configuration (sshd, fail2ban, nftables, systemd unit, SELinux) has a concrete file/syntax above — use them as verifiable templates on a real system.
My notes
Personal notes: points I previously misunderstood, areas I'm still exploring, or lessons from hands-on practice — updated over time.
- The runbook is the thing most worth writing that I put off the longest. When I first took on server duty, every time a Grafana panel turned red I would google
iostat's flags all over again. Eventually I sat down and wrote a proper runbook — "which panel is red → paste which block of commands" (essentially the skeleton of section 2.13) — and I use it almost daily now. The lesson: diagnostic knowledge is only worth anything when it exists as copy-pasteable commands at 2 a.m., not as "I remember there's a tool for that". - I used to misread "high RAM used = about to die". Once I saw a machine at ~90% RAM used, panicked, and was about to propose a RAM upgrade. Only later did I understand that apps (JVM/search-engine types) holding a big heap is designed behavior, and Linux puts idle RAM to work as cache. Since learning about PSI (
/proc/pressure/memory) I barely look at % used anymore: PSI at zero → ignore used entirely; PSI positive and sustained → time to act. Changing one monitored number cut the false alarms dramatically. - The first time I saw
%stealI thought the machine was broken. A cloud VM got inexplicably slow —%usrwas low yet everything dragged. It turned out%stealwas in the double digits: the hypervisor was sharing our CPU with a noisy neighbor. There was nothing to "fix" from inside the VM; the fixable thing was learning to read the right column so I'd stop wasting an afternoon blaming the app. - Disk filling up from accumulated Docker images. On any Docker host, the most common culprit behind a full disk is
/var/lib/docker: every build or deploy stacks on a new image while the old ones are never cleaned up. The fix is a scheduleddocker image prunewith a time filter — and absolutely no casualdocker system pruneon machines with data volumes; one mistaken--volumesflag and real data is gone. A related trap:rm-ing a huge log file withoutdfbudging — the file isn't really dead while a process still holds its fd, andlsof +L1names the culprit. - On firewalls, the biggest trick played on me was iptables-nft.
iptables -Lon one machine andnft list ruleseton another gave me two different pictures and confused me for a good while, until I understood thatiptablesis now usually just a shim translating to nftables (iptables -Vprints(nf_tables)). These days I default tonft list rulesetfor a complete audit, including rules inserted by Docker or fail2ban. - Still exploring: wiring alerts from the SIEM to the firewall so IPs get added to an nftables set automatically past a threshold (fail2ban-style, but fed by centralized logs instead of local ones) — still experimental; I don't dare enable auto-blocking in production yet for fear of banning a NAT range full of real users.