Ansible for beginners: automate your server hardening checklist
13 min read
Your first Ansible playbook: the server hardening checklist, written once and applied to every server the same way. Inventory, modules, handlers, check mode, and verifying the result.

The hardening checklist in the previous article takes about an hour on a new server. Doing it once is fine. The trouble starts with the second server, and the third. One gets a step skipped because you were in a hurry. Another gets a setting changed while debugging and never changed back. Six months later, nobody can say how the servers differ, only that they do.
A checklist done by hand describes what you meant to do. Ansible turns it into a file that describes the state each server should be in, and a tool that connects to the servers and makes them match. Run it on one server or fifty, today or next year, and they end up the same. This article turns the checklist into an Ansible playbook, step by step, and explains each concept as it comes up. The playbooks were tested end to end on Ubuntu 24.04 and Debian 12.
Four ideas before you start
- Ansible runs from your computer, over SSH. Nothing has to be installed on the servers except Python, which Ubuntu and Debian already have.
- The inventory is the list of servers, grouped by role: web servers, databases, and so on.
- Modules do the work. Each one manages one kind of thing: a package, a user, a file, a firewall rule. You say what state you want ("this package installed", "this file with this content"), and the module works out what has to change.
- A playbook is a YAML file with one or more plays. A play names the servers and the tasks to run on them; each task calls a module.
Because modules describe a state rather than a command, running a playbook twice is safe. The second run finds everything already as it should be and changes nothing. This property is called idempotence, and it's what makes Ansible different from a shell script: a script runs its commands again, a playbook checks first.
Install it
Ansible is a Python application. The simplest clean install is with pipx, which keeps it in its own environment:
pipx install --include-deps ansible
ansible --versionModules come in collections, packages grouped by topic, such as community.general, which has the ufw module. The ansible package includes the core tool and the collections used below. The smaller ansible-core package has only the built-in modules.
Tell it about your servers
Create a folder for the project, and in it an inventory file:
# inventory.ini
[web]
web1 ansible_host=203.0.113.10
[web:vars]
ansible_user=deployweb is a group with one server in it, web1, at that IP address. ansible_user is the user Ansible logs in as. It's deploy, the user the playbook will create, because after hardening root can't log in any more.
A small ansible.cfg next to it tells Ansible where the inventory is, so -i inventory.ini becomes optional. The commands below keep it anyway, so they work either way:
# ansible.cfg
[defaults]
inventory = inventory.iniBefore anything else, check that Ansible can reach the server. The server is new and deploy doesn't exist yet, so connect as root this once:
ansible web -i inventory.ini -m ansible.builtin.ping -e ansible_user=rootWhy -e ansible_user=root and not -u root, which looks like the obvious way? Because -u loses to the inventory. Ansible takes settings from many places and ranks them, and a variable set in the inventory, like our ansible_user=deploy, ranks above the -u option. Extra variables given with -e rank above everything, so they always win. The first time, Ansible also asks you to confirm the server's host key, the same way ssh does.
ping here isn't a network ping: it logs in over SSH, runs a bit of Python, and answers pong. If it does, everything Ansible needs is in place.
The checklist as a playbook
Here is the checklist from the previous article, as one playbook. Read it top to bottom; each task is one step, and the names say what they do:
# harden.yml
- name: Harden a new server
hosts: web
become: true
vars:
admin_user: deploy
admin_key: "{{ lookup('ansible.builtin.file', '~/.ssh/id_ed25519.pub') }}"
tasks:
- name: Install all updates
ansible.builtin.apt:
update_cache: true
upgrade: dist
- name: Install the packages the checklist uses
ansible.builtin.apt:
name: [ufw, unattended-upgrades, fail2ban, python3-systemd]
state: present
- name: Create the admin user
ansible.builtin.user:
name: "{{ admin_user }}"
shell: /bin/bash
- name: Let it log in with your key
ansible.posix.authorized_key:
user: "{{ admin_user }}"
key: "{{ admin_key }}"
- name: Let it use sudo without a password
ansible.builtin.copy:
dest: "/etc/sudoers.d/90-{{ admin_user }}"
content: "{{ admin_user }} ALL=(ALL) NOPASSWD:ALL\n"
mode: "0440"
validate: visudo -cf %s
- name: Turn off passwords and root login
ansible.builtin.copy:
dest: /etc/ssh/sshd_config.d/00-hardening.conf
content: |
PermitRootLogin no
PasswordAuthentication no
KbdInteractiveAuthentication no
AllowUsers {{ admin_user }}
mode: "0644"
notify: Reload SSH
- name: Check the whole SSH configuration before it's reloaded
ansible.builtin.command: sshd -t
changed_when: false
check_mode: false
- name: Allow SSH through the firewall
community.general.ufw:
rule: allow
name: OpenSSH
- name: Allow web traffic
community.general.ufw:
rule: allow
port: "{{ item }}"
proto: tcp
loop: ["80", "443"]
- name: Deny all other incoming traffic
community.general.ufw:
default: deny
direction: incoming
- name: Turn the firewall on
community.general.ufw:
state: enabled
- name: Install security updates every day
ansible.builtin.copy:
dest: /etc/apt/apt.conf.d/20auto-upgrades
content: |
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Unattended-Upgrade "1";
mode: "0644"
- name: Reboot at 04:00 when an update needs it
ansible.builtin.copy:
dest: /etc/apt/apt.conf.d/52unattended-upgrades-local
content: |
Unattended-Upgrade::Automatic-Reboot "true";
Unattended-Upgrade::Automatic-Reboot-Time "04:00";
mode: "0644"
- name: Block repeated SSH login attempts
ansible.builtin.copy:
dest: /etc/fail2ban/jail.local
content: |
[sshd]
enabled = true
backend = systemd
maxretry = 5
findtime = 10m
bantime = 1h
mode: "0644"
notify: Restart fail2ban
handlers:
- name: Reload SSH
ansible.builtin.service:
name: ssh
state: reloaded
- name: Restart fail2ban
ansible.builtin.service:
name: fail2ban
state: restartedA few things in it are worth a closer look, because they're the core of how Ansible works:
become: true runs the tasks with sudo. Ansible logs in as an ordinary user and becomes root only for the work, the same way you would by hand.
Modules, not commands. Almost every task uses a module that knows its subject: apt for packages, user for users, copy for files with a given content, ufw for the firewall. That's what makes the playbook safe to run again.
copy compares the file on the server with the content in the playbook and writes it only if they differ. A command task, on the other hand, runs every time. Ansible can't know what it changed, so it reports changed by default. The one in this playbook (sshd -t) only checks, so it's marked changed_when: false. Its check_mode: false comes up again below.
validate checks a file before it replaces the old one. For the sudoers file this matters a lot: a syntax error there can break sudo for everyone, and visudo -cf refuses the file before that can happen.
Handlers run once, at the end of the play, and only if a task that notifies them changed something. If the SSH settings are already in place, SSH isn't reloaded.
Because handlers run last, the sshd -t check always comes before the reload. If the configuration is broken, the play stops with an error, and the running SSH server is never told to read it. The broken file is still on disk, though, and the next restart or reboot would read it. Fix it and run the playbook again straight away.
Variables such as admin_user are written once at the top and used everywhere with {{ }}. Changing the username means changing one line.
The order of the tasks matters as much as in the checklist. The new user gets its key and its sudo before root's login is turned off. The firewall allows SSH before it's turned on.
On the passwordless sudo: the deploy user can only log in with your key, and Ubuntu's own cloud images give their default user the same rule. If you'd rather have a password, drop the sudoers task, set a password for the user, and run playbooks with --ask-become-pass.
Run it: first as root, then as yourself
The first run has to log in as root, because deploy doesn't exist yet. As with ping, that means -e, not -u:
ansible-playbook -i inventory.ini harden.yml -e ansible_user=rootAnsible prints each task with ok (already as it should be) or changed (it changed something), and a summary at the end. On a new server most tasks are changed. If the updates included a new kernel, restart the server once so it runs: ansible web -i inventory.ini -m ansible.builtin.reboot -b. The reboot module waits until the server is back.
From now on, root can't log in, and Ansible uses deploy from the inventory. Run the playbook again:
ansible-playbook -i inventory.ini harden.ymlThis time everything should be ok and the summary should say changed=0. That second run is the test that the playbook is idempotent. A task that reports changed on every run is either doing work it doesn't need to, or not describing a state at all, and it's worth fixing.
Before you change a playbook on a server that's already running, ask Ansible what it would do without doing it:
ansible-playbook -i inventory.ini harden.yml --check --diff--check runs in dry-run mode, and --diff shows the changes each file would get, line by line. Modules that can't predict their effect, such as command, are skipped in check mode. That's why the sshd -t task has check_mode: false: it only reads, so it's safe to run for real even in a dry run, and the dry run then checks the SSH configuration too.
For a new server, add it to the inventory and run the same playbook, the first time as root with -e ansible_user=root and --limit set to the new server. It gets exactly what the others have.
A playbook that finished isn't a server that works
When ansible-playbook ends without an error, it has proved that Ansible ran. It hasn't proved that the server does what you wanted. The previous article showed how SSH can keep accepting passwords while its configuration says it doesn't. A playbook that copies the right file would report success in exactly that case.
So add a second play that checks the result, the same way you would by hand, and make it fail loudly when something is wrong:
# verify.yml
- name: Verify the server does what the checklist says
hosts: web
become: true
gather_facts: false
tasks:
- name: Read the settings SSH will actually use
ansible.builtin.command: sshd -T
register: sshd_effective
changed_when: false
check_mode: false
- name: Passwords and root login are off
ansible.builtin.assert:
that:
- "'passwordauthentication no' in sshd_effective.stdout_lines"
- "'kbdinteractiveauthentication no' in sshd_effective.stdout_lines"
- "'permitrootlogin no' in sshd_effective.stdout_lines"
fail_msg: SSH still accepts passwords or root. Check the other files in /etc/ssh/sshd_config.d/.
- name: Read the firewall status
ansible.builtin.command: ufw status verbose
register: ufw_status
changed_when: false
check_mode: false
- name: The firewall is on and denies by default
ansible.builtin.assert:
that:
- "'Status: active' in ufw_status.stdout"
- "'deny (incoming)' in ufw_status.stdout"
fail_msg: The firewall is off or allows incoming traffic by default.register saves a task's output in a variable, and assert fails the play if a condition isn't true. Run it after the hardening, and again whenever you want to know whether a server has drifted: ansible-playbook -i inventory.ini verify.yml.
It catches exactly the trap from the previous article. On a test server with a 50-cloud-init.conf that turned passwords on, I renamed the hardening file to 60-hardening.conf. SSH accepted passwords again, and this play failed with the message above. The hardening playbook alone would still have reported success.
I learned to do this the hard way, in a Kubernetes lab I built with Terraform and Ansible. The playbook finished without an error, all four nodes reported Ready, and nothing could reach a service on another node. The network layer had picked the public interface instead of the private one, and the firewall dropped the traffic between the nodes there. Every check I had looked at said the cluster was fine.
Since then, the playbook there ends with a verify play that asserts the things that once broke while everything visible said otherwise, including a real request between nodes.
Good practices as it grows
The playbook above is deliberately one file, with one server in a text file. That's the right way to start. These are the practices that keep it manageable as the number of servers and tasks grows.
Let the inventory come from the provider. A static inventory has to be edited every time a server is created or removed, and sooner or later it's wrong. A dynamic inventory asks the cloud provider's API for the current list instead. Ansible has inventory plugins for AWS, Azure, Google Cloud, Hetzner and others. For Hetzner Cloud it's a small YAML file, whose name has to end in hcloud.yml:
# inventory/hcloud.yml
plugin: hetzner.hcloud.hcloud
# The API token is read from the HCLOUD_TOKEN environment variable
label_selector: env=production
keyed_groups:
# A server labelled role=web lands in the group role_web
- key: hcloud_labels.role
prefix: roleThe plugin needs two Python libraries next to Ansible. With the install above, add them with pipx inject ansible requests python-dateutil.
Servers are selected by label and grouped by label, so a new server labelled role=web is in the right group the moment it exists, with nothing to edit. ansible-inventory --graph shows the groups the plugin produced. With this file, the playbook's hosts: web becomes hosts: role_web.
When Terraform creates the servers, the labels come from the same code, and the two can't drift apart. That's how my Kubernetes lab works: Terraform creates and labels the servers, and Ansible finds them by those labels.
Move variables out of the inventory. Make inventory a directory, and put variables in group_vars/web.yml and host_vars/web1.yml next to the inventory file or plugin. Ansible picks them up by name, and they work the same with a static or a dynamic inventory. The inventory stays a list of servers, and each group's settings live in one file.
Split it into roles. A role is a folder with its own tasks, handlers, files and default variables: ssh, firewall, updates. A playbook then lists the roles each group of servers gets, and the same role serves every playbook that needs it.
Use existing roles, but read them first. Ansible Galaxy has roles and collections for almost everything, including complete hardening roles. They're a good starting point, but a hardening role changes many settings at once, and some of them may not suit your servers. Read and adapt them before you run them.
Pin every version. List the collections in a requirements.yml with exact versions, install them into the project (collections_path = ./collections in ansible.cfg), and pin Ansible itself too, with exact versions in a requirements.txt or a lock file from a tool such as uv. Otherwise the same playbook can behave differently on a colleague's computer, or next year on yours.
Change running servers carefully. --limit web1 runs a playbook on one server first. serial: 1 in a play changes servers one at a time, and stops if one fails, so a bad change never reaches all of them at once. Tags (tags: [ssh] on tasks, --tags ssh on the command line) run only the part you changed.
Test the playbooks, not just the servers. ansible-playbook --syntax-check catches broken YAML and misplaced keywords, but not a misspelled module option: that only fails when the task runs. ansible-lint checks module options too, along with common mistakes such as a command task with no changed_when that reports a change on every run. Run both in CI on every change, so a broken playbook never reaches a server. For roles, Molecule creates a throwaway server, applies the role, runs it a second time to check that nothing changes, and destroys the server again.
Keep secrets out of plain text. Ansible Vault encrypts passwords and API keys, so they can live in the same repository as the playbooks. Mark the tasks that use them no_log: true, so they don't end up in the output or in CI logs.
Make it fast. Ansible already reuses SSH connections between tasks, through SSH's ControlPersist. Pipelining is off by default: pipelining = True under [ssh_connection] in ansible.cfg cuts the SSH operations per task further, and on many servers or long playbooks the difference is noticeable. It needs sudo without requiretty, which is the default on Ubuntu and Debian.
On my own servers I started from existing Ansible hardening roles and adapted them, then audited the result with Lynis and OpenSCAP and fixed what they found. The audit is the same idea as the verify play, on a larger scale: don't trust that the configuration was applied, check what the server actually does.
The checklist from the previous article is now a file you can read, review, keep in Git and run on the next server in minutes, with the same result every time. That's the real gain: not the hour saved on one server, but knowing that every server is the same, and being able to prove it. If you're moving servers from hand-made to automated and want a second pair of eyes, tell me what you're working on.
Sources: Installing Ansible; Intro to playbooks; Handlers; Check mode and diff mode; Variable precedence; Dynamic inventory; hetzner.hcloud inventory plugin; Molecule; community.general.ufw module; ansible.posix.authorized_key module; Ansible Lint.


