Compare commits

..

6 Commits

Author SHA1 Message Date
Timothy Stewart e0ac53dba3 fix(metallb): verify the metallb-system namespace actually exists
- change the Test metallb-system namespace task to run
  k3s kubectl get namespace metallb-system instead of the bare
  -n metallb-system, which printed kubectl usage and always exited 0
- add a regression test that asserts the task uses an explicit get
  and fails if it ever regresses to the usage-only form
- wire the new test into pre-commit
2026-08-06 09:08:11 -05:00
Techno Tim 287d8b7a27 feat(kube-vip): add endpoint override for the internal listening address (#699)
* feat(kube-vip): add endpoint override for the internal listening address

Add a kube_vip_endpoint variable so the address kube-vip binds and listens on
can differ from the announced apiserver_endpoint. This is useful for complex
routing and site-to-site tunnels where the VIP kube-vip advertises over ARP
differs from the address it listens on internally.

- roles/k3s_server/templates/vip.yaml.j2: use kube_vip_endpoint (defaulting to
  apiserver_endpoint) for the `address` env and for deriving `vip_subnet`
- roles/k3s_server/defaults/main.yml: add kube_vip_endpoint default (null)
- roles/k3s_server/meta/main.yml: add kube_vip_endpoint argument_spec
- inventory/sample/group_vars/all.yml: document the new sample variable
- README.md: document the kube_vip_endpoint option
- .github/scripts/test-kube-vip-manifest.py: extend regression test to cover the
  default (apiserver_endpoint) and the override case

Closes #221

* chore(ci): extend molecule job timeout to 3 hours

The default scenario occasionally takes longer than 150 minutes on the shared
nested-virt runner (k3s agent notify-wait can exceed the limit under load), and
a single timeout aborts the whole run before the other four scenarios execute.
Raise timeout-minutes from 150 to 180 so a slow-but-progressing run completes
instead of aborting.

The default scenario remains first in the matrix so a failure surfaces fastest.

* fix(kube-vip): fall back on null kube_vip_endpoint and cover it in the test

- vip_subnet and address use default(apiserver_endpoint, true) so the null
  role default falls back to the apiserver endpoint instead of rendering an
  empty/invalid address and subnet
- change the manifest regression test default case to pass kube_vip_endpoint
  as None so it pins the real runtime null condition and fails fast on this
  regression rather than timing out in CI
2026-08-06 02:48:57 -05:00
Techno Tim 56bb912bd3 feat(prereq): add toggle to disable swap on all cluster nodes (#698)
k3s recommends swap be disabled on every node. Add a disable_swap toggle
(default true) to the prereq role that turns swap off now and comments out the
swap entries in /etc/fstab so swap stays off across reboots. The change is
applied uniformly across all k3s_cluster hosts (all-or-nothing) since leaving
swap on for some nodes but not others creates uneven scheduling and latency.

- roles/prereq/defaults/main.yml: add disable_swap: true default
- roles/prereq/tasks/main.yml: idempotent swapoff -a and fstab comment-out block
  gated on disable_swap
- inventory/sample/group_vars/all.yml: document the new sample variable
- README.md: document the disable_swap option
- .github/scripts/test-disable-swap.sh: regression test
- .pre-commit-config.yaml: wire the test into pre-commit

Closes #670
2026-08-05 07:39:00 -05:00
Techno Tim cf76292169 feat(site): fail fast when cluster nodes share a hostname (#697)
Add a preflight assertion to site.yml that confirms every host in the
k3s_cluster group reports a unique hostname. k3s registers each node keyed by
its hostname, so duplicate hostnames prevent nodes from joining the cluster
and are hard to troubleshoot. The check fails fast in the Pre tasks instead of
surfacing as a cryptic registration failure later.

- site.yml: assert (cluster_hostnames | unique | length) == length for the
  k3s_cluster group, skipping hosts outside that group
- README.md: document the unique-hostname requirement
- .github/scripts/test-unique-hostname-precheck.sh: regression test
- .pre-commit-config.yaml: wire the test into pre-commit

Closes #636
2026-08-05 05:14:29 -05:00
Techno Tim 52c086d638 feat(metallb): add option to limit MetalLB layer2 announcements to specific interfaces (#696)
- roles/k3s_server_post/templates/metallb.crs.j2: render spec.interfaces in the
  L2Advertisement when metal_lb_interfaces is a non-empty list, so MetalLB only
  announces on the configured interfaces. Empty list (default) keeps announcing
  on all interfaces, preserving existing behavior.
- roles/k3s_server_post/defaults/main.yml: add metal_lb_interfaces: [] default
- roles/k3s_server_post/meta/main.yml: add metal_lb_interfaces argument_spec
- inventory/sample/group_vars/all.yml: document the new sample variable
- .github/scripts/test-metallb-interfaces.py: regression test rendering the
  template for empty/single/multiple interfaces and confirming the BGP path is
  unaffected
- .pre-commit-config.yaml: wire the new test into pre-commit

Co-authored-by: Leo <leo@kuboschek.me>
2026-08-05 02:48:24 -05:00
Techno Tim 010551b8d2 feat(cilium): add toggle to enable or disable the Envoy proxy (#695)
Add a `cilium_envoy` variable (default true, matching upstream Cilium 1.20
which installs Envoy by default) that controls whether the Envoy proxy is
deployed for Cilium L7 policies. Pass it to Helm as `envoy.enabled` so users
with no L7 policies can skip Envoy to save resources.

- roles/k3s_server_post/defaults/main.yml: add cilium_envoy: true default
- roles/k3s_server_post/tasks/cilium.yml: add --helm-set envoy.enabled to the
  install/upgrade command, driven by the cilium_envoy conditional
- inventory/sample/group_vars/all.yml: document cilium_envoy sample var
- .github/scripts/test-cilium-envoy-toggle.py: regression test asserting the
  install command carries the envoy.enabled helm-set and renders true/false
- .pre-commit-config.yaml: wire the new test into pre-commit

Co-authored-by: Léo Nonnenmacher <leo@nonnenmacher-logel.fr>
2026-08-05 00:00:39 -05:00
21 changed files with 510 additions and 4 deletions
+120
View File
@@ -0,0 +1,120 @@
#!/usr/bin/env python3
"""Regression test for the Cilium Envoy toggle.
The `cilium_envoy` variable lets users enable or disable the Cilium Envoy
proxy. The Install/upgrade Cilium task in
roles/k3s_server_post/tasks/cilium.yml passes the value through to Helm as
`envoy.enabled`. This test:
- loads the real "Install Cilium" task and confirms the install/upgrade
command actually contains the `envoy.enabled` Helm value,
- renders the conditional that computes the Helm value and confirms it
produces `true` when cilium_envoy is enabled and `false` when disabled,
- confirms the task stays forward/backward compatible (no raw `true` /
`false` hardcoded in place of the conditional).
"""
from __future__ import print_function
import os
import re
import subprocess
import yaml
from jinja2 import Environment
ENVOY_EXPRESSION = '{{ "true" if cilium_envoy else "false" }}'
def repo_root():
return subprocess.check_output(
["git", "rev-parse", "--show-toplevel"], text=True
).strip()
def fail(message):
raise SystemExit("Cilium Envoy toggle test failed: " + message)
def extract_install_command(path):
"""Return the command string for the 'Install Cilium' task.
Walks both top-level tasks and tasks nested inside a `block`/`always`/
`rescue` list, since the Cilium deploy steps are grouped under the
'Prepare Cilium CLI on first master and deploy CNI' block.
"""
with open(path, encoding="utf-8") as handle:
doc = yaml.safe_load(handle)
def find_command(tasks):
for task in tasks:
if not isinstance(task, dict):
continue
if task.get("name") == "Install Cilium":
command = task.get("ansible.builtin.command")
if command is None:
raise SystemExit(
"Cilium Envoy toggle test failed: "
"'Install Cilium' task has no ansible.builtin.command"
)
return command
# Recurse into block/always/rescue sub-lists.
for key in ("block", "always", "rescue"):
nested = task.get(key)
if isinstance(nested, list):
found = find_command(nested)
if found is not None:
return found
return None
command = find_command(doc)
if command is None:
raise SystemExit(
"Cilium Envoy toggle test failed: could not find 'Install Cilium' task"
)
return command
def assert_envoy_in_command(command):
if "envoy.enabled" not in command:
fail("install command is missing --helm-set envoy.enabled")
if ENVOY_EXPRESSION not in command:
fail(
"install command does not use the cilium_envoy conditional: "
"expected {0!r}".format(ENVOY_EXPRESSION)
)
# The conditional must be a WYSIWYG helm-set value, not a pre-rendered
# true/false literal (which would ignore the cilium_envoy variable).
if re.search(r"--helm-set envoy\.enabled=true(?:$|\s)", command):
fail("install command hardcodes envoy.enabled=true")
if re.search(r"--helm-set envoy\.enabled=false(?:$|\s)", command):
fail("install command hardcodes envoy.enabled=false")
def assert_render():
env = Environment()
def render_for(value):
template = env.from_string(ENVOY_EXPRESSION)
return template.render(cilium_envoy=value)
if render_for(True) != "true":
fail("envoy conditional did not render 'true' when enabled")
if render_for(False) != "false":
fail("envoy conditional did not render 'false' when disabled")
def main():
root = repo_root()
cilium_tasks = os.path.join(
root, "roles", "k3s_server_post", "tasks", "cilium.yml"
)
command = extract_install_command(cilium_tasks)
assert_envoy_in_command(command)
assert_render()
print("Cilium Envoy toggle regression test passed")
if __name__ == "__main__":
main()
+37
View File
@@ -0,0 +1,37 @@
#!/usr/bin/env bash
set -Eeuo pipefail
repo_root="$(git rev-parse --show-toplevel)"
prereq_defaults="$repo_root/roles/prereq/defaults/main.yml"
prereq_tasks="$repo_root/roles/prereq/tasks/main.yml"
# #670: k3s recommends swap be disabled on all nodes. The prereq role must expose
# a disable_swap toggle (defaulting to true) that turns swap off now and comments
# out the /etc/fstab swap entries so swap stays off across reboots.
grep -Eq -- '^disable_swap: true' "$prereq_defaults" || {
printf 'prereq defaults are missing disable_swap: true\n' >&2
exit 1
}
grep -Fq -- 'Disable swap on all cluster nodes' "$prereq_tasks" || {
printf 'prereq tasks are missing the swap-disable block\n' >&2
exit 1
}
grep -Fq -- 'swapoff -a' "$prereq_tasks" || {
printf 'swap-disable block does not run swapoff -a\n' >&2
exit 1
}
grep -Fq -- '/etc/fstab' "$prereq_tasks" || {
printf 'swap-disable block does not comment out /etc/fstab swap entries\n' >&2
exit 1
}
if ! grep -Eq -- 'when: disable_swap' "$prereq_tasks"; then
printf 'swap-disable block is not gated on the disable_swap toggle\n' >&2
exit 1
fi
printf 'Swap disable regression test passed\n'
+32
View File
@@ -123,6 +123,38 @@ def main():
if "name: bgp_peers" in output:
fail("bgp_peers present even though the peer list is empty")
# kube_vip_endpoint defaults to null (defined in role defaults): the
# address and subnet must fall back to the apiserver endpoint. default()
# without a truthy flag does NOT fall back on null, only on undefined, so
# this case pins the null runtime condition to prevent that regression.
output = render(
env,
{
"_kube_vip_bgp_peers": [],
"kube_vip_endpoint": None,
"kube_vip_arp": True,
"kube_vip_bgp": False,
},
)
if "value: 192.168.30.222" not in output:
fail("null kube_vip_endpoint does not fall back to apiserver_endpoint")
# kube_vip_endpoint set: overrides the internal listening address AND the
# subnet derivation while the advertised apiserver_endpoint stays separate.
output = render(
env,
{
"_kube_vip_bgp_peers": [],
"kube_vip_endpoint": "10.66.1.5",
"kube_vip_arp": True,
"kube_vip_bgp": False,
},
)
if "value: 10.66.1.5" not in output:
fail("kube_vip_endpoint did not override the address")
if "value: 192.168.30.222" in output:
fail("apiserver_endpoint leaked into address when kube_vip_endpoint set")
print("kube-vip manifest regression test passed")
@@ -0,0 +1,88 @@
#!/usr/bin/env python3
"""Regression test for the MetalLB L2Advertisement interfaces.
`metal_lb_interfaces` restricts which network interfaces MetalLB announces
load balancer IPs on in layer2 mode. When the list is non-empty, the
L2Advertisement in roles/k3s_server_post/templates/metallb.crs.j2 must render
a `spec.interfaces` block; when it is empty (the default), no spec is rendered
so MetalLB announces on all interfaces.
This renders the template and asserts both cases plus the BGP path (which must
not be affected by the L2 interfaces variable).
"""
from __future__ import print_function
import os
import subprocess
from jinja2 import Environment, FileSystemLoader, StrictUndefined
def repo_root():
return subprocess.check_output(
["git", "rev-parse", "--show-toplevel"], text=True
).strip()
def fail(message):
raise SystemExit("MetalLB interfaces test failed: " + message)
def render(env, extra_vars):
base_vars = {
"metal_lb_mode": "layer2",
"metal_lb_ip_range": "192.168.30.80-192.168.30.90",
}
base_vars.update(extra_vars)
template = env.get_template("metallb.crs.j2")
return template.render(**base_vars)
def main():
root = repo_root()
template_dir = os.path.join(
root, "roles", "k3s_server_post", "templates"
)
env = Environment(
loader=FileSystemLoader(template_dir), undefined=StrictUndefined
)
# Empty list (default): no spec.interfaces in the L2Advertisement.
output = render(env, {"metal_lb_interfaces": []})
if "spec:\n interfaces:" in output:
fail("spec.interfaces rendered with an empty metal_lb_interfaces")
if "kind: L2Advertisement" not in output:
fail("L2Advertisement missing in layer2 mode")
# Single interface.
output = render(env, {"metal_lb_interfaces": ["eth1"]})
if "spec:\n interfaces:\n - eth1" not in output:
fail("single interface was not rendered in spec.interfaces")
# Multiple interfaces.
output = render(env, {"metal_lb_interfaces": ["eth1", "eth2"]})
if "spec:\n interfaces:\n - eth1\n - eth2" not in output:
fail("multiple interfaces were not rendered in spec.interfaces")
# BGP mode must not emit an L2Advertisement spec at all.
output = render(
env,
{
"metal_lb_mode": "bgp",
"metal_lb_interfaces": ["eth1"],
"metal_lb_bgp_my_asn": "64513",
"metal_lb_bgp_peer_asn": "64512",
"metal_lb_bgp_peer_address": "192.168.30.1",
},
)
if "kind: L2Advertisement" in output:
fail("L2Advertisement rendered in bgp mode")
if "interfaces:" in output:
fail("interfaces rendered in bgp mode")
print("MetalLB interfaces regression test passed")
if __name__ == "__main__":
main()
+72
View File
@@ -0,0 +1,72 @@
#!/usr/bin/env python3
"""Regression test for the MetalLB namespace existence check.
The "Test metallb-system namespace" task in
roles/k3s_server_post/tasks/metallb.yml must actually verify the namespace
exists. A previous version ran `k3s kubectl -n metallb-system` with no
subcommand, which only printed a usage page and always exited 0, so the task
always succeeded even when the namespace did not exist (issue #350).
This test loads the real task and asserts the command performs an explicit
`get namespace metallb-system`, which returns non-zero when the namespace is
absent.
"""
from __future__ import print_function
import os
import subprocess
import yaml
def repo_root():
return subprocess.check_output(
["git", "rev-parse", "--show-toplevel"], text=True
).strip()
def fail(message):
raise SystemExit(
"MetalLB namespace test failed: " + message
)
def main():
task_file = os.path.join(
repo_root(), "roles", "k3s_server_post", "tasks", "metallb.yml"
)
with open(task_file, encoding="utf-8") as handle:
tasks = yaml.safe_load(handle)
task = None
for entry in tasks:
if entry.get("name") == "Test metallb-system namespace":
task = entry
break
if task is None:
fail("could not find the 'Test metallb-system namespace' task")
cmd = task.get("ansible.builtin.command")
if not cmd:
cmd = task.get("command")
if not cmd:
fail("task does not use ansible.builtin.command")
command_text = cmd if isinstance(cmd, str) else " ".join(cmd)
# A bare `-n metallb-system` with no subcommand prints kubectl usage and
# always exits 0, so it never proves the namespace exists. The fix must
# use an explicit get.
if "get namespace metallb-system" not in command_text:
fail(
"command does not run 'get namespace metallb-system'; "
"the task would only print usage and never verify the namespace "
"(got: {0!r})".format(command_text)
)
print("MetalLB namespace check regression test passed")
if __name__ == "__main__":
main()
+28
View File
@@ -0,0 +1,28 @@
#!/usr/bin/env bash
set -Eeuo pipefail
repo_root="$(git rev-parse --show-toplevel)"
site_play="$repo_root/site.yml"
# #636: verify the "Pre tasks" play asserts that all k3s_cluster hosts have
# unique hostnames, so a duplicate-hostname inventory fails fast instead of
# silently breaking node registration/joining.
grep -Fq -- 'Verify all cluster nodes have unique hostnames' "$site_play" || {
printf 'site.yml is missing the unique-hostname preflight check\n' >&2
exit 1
}
# The check must deduplicate the cluster hostname list via the `unique` filter
# and compare lengths, i.e. groups['k3s_cluster'] must be referenced.
grep -Fq -- "groups['k3s_cluster']" "$site_play" || {
printf 'unique-hostname check does not iterate the k3s_cluster group\n' >&2
exit 1
}
if ! grep -Eq -- 'cluster_hostnames.*\|.*unique|\| unique' "$site_play"; then
printf 'unique-hostname check does not deduplicate the hostname list\n' >&2
exit 1
fi
printf 'Unique hostname preflight regression test passed\n'
+1 -1
View File
@@ -88,7 +88,7 @@ jobs:
trap stop_monitor EXIT
/usr/bin/time -v -o "$timing_file" \
molecule test --scenario-name ${{ matrix.scenario }}
timeout-minutes: 150
timeout-minutes: 180
env:
ANSIBLE_K3S_LOG_DIR: ${{ runner.temp }}/logs/k3s-ansible/${{ matrix.scenario }}
ANSIBLE_SSH_RETRIES: 4
+37
View File
@@ -62,6 +62,18 @@ repos:
language: system
pass_filenames: false
files: ^roles/k3s_server/tasks/(main|join_master)\.yml$|^\.github/scripts/test-k3s-server-bootstrap\.sh$
- id: unique-hostname-precheck-test
name: Unique hostname precheck test
entry: .github/scripts/test-unique-hostname-precheck.sh
language: system
pass_filenames: false
files: ^site\.yml$|^\.github/scripts/test-unique-hostname-precheck\.sh$
- id: disable-swap-test
name: Disable swap test
entry: .github/scripts/test-disable-swap.sh
language: system
pass_filenames: false
files: ^roles/prereq/(tasks/main|defaults/main)\.yml$|^\.github/scripts/test-disable-swap\.sh$
- id: cilium-bgp-manifest-test
name: Cilium BGP manifest test
entry: python3 .github/scripts/test-cilium-bgp-manifest.py
@@ -70,6 +82,15 @@ repos:
- Jinja2>=3.1
pass_filenames: false
files: ^roles/k3s_server_post/templates/cilium\.crs\.j2$|^\.github/scripts/test-cilium-bgp-manifest\.py$
- id: cilium-envoy-toggle-test
name: Cilium Envoy toggle test
entry: python3 .github/scripts/test-cilium-envoy-toggle.py
language: python
additional_dependencies:
- Jinja2>=3.1
- PyYAML
pass_filenames: false
files: ^roles/k3s_server_post/tasks/cilium\.yml$|^\.github/scripts/test-cilium-envoy-toggle\.py$
- id: kube-vip-manifest-test
name: kube-vip manifest test
entry: python3 .github/scripts/test-kube-vip-manifest.py
@@ -84,6 +105,22 @@ repos:
language: system
pass_filenames: false
files: ^roles/k3s_server/tasks/metallb\.yml$|^\.github/scripts/test-metallb-remote-read\.sh$
- id: metallb-interfaces-test
name: MetalLB interfaces test
entry: python3 .github/scripts/test-metallb-interfaces.py
language: python
additional_dependencies:
- Jinja2>=3.1
pass_filenames: false
files: ^roles/k3s_server_post/templates/metallb\.crs\.j2$|^\.github/scripts/test-metallb-interfaces\.py$
- id: metallb-namespace-test
name: MetalLB namespace test
entry: python3 .github/scripts/test-metallb-namespace.py
language: python
additional_dependencies:
- PyYAML
pass_filenames: false
files: ^roles/k3s_server_post/tasks/metallb\.yml$|^\.github/scripts/test-metallb-namespace\.py$
- id: metallb-deploy-condition-test
name: MetalLB deploy condition test
entry: python3 .github/scripts/test-metallb-deploy-condition.py
+6
View File
@@ -50,6 +50,10 @@ Supported processor architectures are:
- Server and agent nodes should support passwordless SSH access. Otherwise, pass `--ask-pass --ask-become-pass` to
each playbook command.
- Every node in the cluster must have a **unique hostname**. k3s registers each node keyed by its hostname, so
two nodes with the same hostname cannot join the cluster. `site.yml` asserts this up front and fails fast if
any duplicate is found.
## 🚀 Getting Started
### 🍴 Preparation
@@ -215,6 +219,7 @@ See the commands [here](https://technotim.com/posts/k3s-etcd-ansible/#testing-yo
| `k3s_server` | `kube_vip_bgp_peers` | list | `[]` | Not required | List of BGP peer ASN & address pairs |
| `k3s_server` | `kube_vip_bgp_peers_groups` | list | `['k3s_master']` | Not required | Inventory group in which to search for additional `kube_vip_bgp_peers` parameters to merge. |
| `k3s_server` | `kube_vip_iface` | string | `~` | Not required | Explicitly define an interface that ALL control nodes should use to propagate the VIP, define it here. Otherwise, kube-vip will determine the right interface automatically at runtime. |
| `k3s_server` | `kube_vip_endpoint` | string | `~` | Not required | Overrides the internal address kube-vip binds/listens on, which can differ from the announced apiserver_endpoint for complex routing/tunnels. Defaults to apiserver_endpoint. |
| `k3s_server` | `kube_vip_tag_version` | string | `v1.2.2` | Not required | Image tag for kube-vip |
| `k3s_server` | `kube_vip_cloud_provider_tag_version` | string | `v0.0.12` | Not required | Tag for kube-vip-cloud-provider manifest when enable |
| `k3s_server`, `k3_server_post` | `kube_vip_lb_ip_range` | string | `~` | Not required | IP range for kube-vip load balancer |
@@ -257,6 +262,7 @@ See the commands [here](https://technotim.com/posts/k3s-etcd-ansible/#testing-yo
| `reboot` (playbook) | `concurrent_reboots` | int/string | `100%` | Not required | Number (or percentage) of nodes to reboot at a time for a staggered reboot |
| `reboot` (playbook) | `wait_seconds_after_reboot` | int | `0` | Not required | Pause in seconds between staggered reboot batches |
| `prereq` | `system_timezone` | string | `null` | Not required | Timezone to be set on all nodes |
| `prereq` | `disable_swap` | bool | `true` | Not required | Disable swap on all cluster nodes (swapoff + comment out /etc/fstab swap entries), all-or-nothing |
| `proxmox_lxc`, `reset_proxmox_lxc` | `proxmox_lxc_ct_ids` | list | ❌ | Required | Proxmox container ID list |
| `raspberrypi` | `state` | string | `present` | Not required | Indicates whether the k3s prerequisites for Raspberry Pi should be set up (possible values are `present` and `absent`) |
+19
View File
@@ -7,6 +7,10 @@ systemd_dir: /etc/systemd/system
# Set your timezone
system_timezone: Your/Timezone
# k3s recommends swap be disabled on every cluster node. Applied uniformly to all
# nodes (all-or-nothing) in the prereq role. Set to false to leave swap enabled.
disable_swap: true
# interface which will be used for flannel
# Defaults to each host's default IPv4 interface (e.g. eth0, enp1s0, ens3)
# so KVM/cloud hosts without eth0 work out of the box. Override per-host if needed.
@@ -24,6 +28,10 @@ cilium_mode: native # native when nodes are on the same subnet or use BGP, other
cilium_tag: v1.20.0 # cilium version tag
cilium_cli_tag: v0.19.7 # cilium cli version tag
cilium_hubble: true # enable hubble observability relay and ui
cilium_envoy: true # enable the Envoy proxy for Cilium L7 policies
# disable cilium_envoy to skip the Envoy proxy entirely (e.g. no L7 policies)
# cilium_envoy: false
# if using calico or cilium, you may specify the cluster pod cidr pool
cluster_cidr: 10.52.0.0/16
@@ -40,6 +48,11 @@ cilium_bgp_lb_cidr: 192.168.31.0/24 # cidr for cilium loadbalancer ipam
# enable kube-vip ARP broadcasts
kube_vip_arp: true
# (optional) overrides the address kube-vip binds/listens on internally, which
# can differ from the announced apiserver_endpoint for complex routing/tunnels.
# Defaults to apiserver_endpoint. Also used to derive the kube-vip subnet.
# kube_vip_endpoint: 10.66.1.5
# enable kube-vip BGP peering
kube_vip_bgp: false
@@ -116,6 +129,12 @@ metal_lb_controller_tag_version: v0.16.0
# metallb ip range for load balancer
metal_lb_ip_range: 192.168.30.80-192.168.30.90
# (optional) limit MetalLB layer2 announcements to specific network interfaces.
# Leave empty (default) to announce on all interfaces.
# metal_lb_interfaces:
# - eth1
# - eth2
# Only enable if your nodes are proxmox LXC nodes, make sure to configure your proxmox nodes
# in your hosts.ini file.
# Please read https://gist.github.com/triangletodd/02f595cd4c0dc9aac5f7763ca2264185 before using this.
+1
View File
@@ -7,6 +7,7 @@ group_name_master: master
kube_vip_arp: true
kube_vip_iface:
kube_vip_endpoint:
kube_vip_cloud_provider_tag_version: v0.0.12
kube_vip_tag_version: v1.2.2
+9
View File
@@ -78,6 +78,15 @@ argument_specs:
- automatically at runtime.
default: ~
kube_vip_endpoint:
description:
- Overrides the address kube-vip binds/listens on internally, which
- can differ from the announced apiserver_endpoint for complex
- routing and site-to-site tunnels.
- Defaults to apiserver_endpoint and is used to derive the kube-vip
- subnet.
default: ~
kube_vip_tag_version:
description: Image tag for kube-vip
default: v1.2.2
+2 -2
View File
@@ -37,7 +37,7 @@ spec:
value: {{ kube_vip_iface }}
{% endif %}
- name: vip_subnet
value: "{{ apiserver_endpoint | ansible.utils.ipsubnet | ansible.utils.ipaddr('prefix') }}"
value: "{{ (kube_vip_endpoint | default(apiserver_endpoint, true)) | ansible.utils.ipsubnet | ansible.utils.ipaddr('prefix') }}"
- name: cp_enable
value: "true"
- name: cp_namespace
@@ -55,7 +55,7 @@ spec:
- name: vip_retryperiod
value: "2"
- name: address
value: {{ apiserver_endpoint }}
value: {{ kube_vip_endpoint | default(apiserver_endpoint, true) }}
{% if kube_vip_bgp | default(false) | bool %}
{% if kube_vip_bgp_routerid is defined %}
- name: bgp_routerid
+2
View File
@@ -18,6 +18,7 @@ cilium_bgp_peer_asn: 64512
cilium_bgp_neighbors: []
cilium_bgp_neighbors_groups: ['k3s_all']
cilium_bgp_lb_cidr: 192.168.31.0/24
cilium_envoy: true
cilium_hubble: true
cilium_mode: native
cilium_tag: v1.20.0
@@ -38,4 +39,5 @@ group_name_master: master
metal_lb_mode: layer2
metal_lb_available_timeout: 240s
metal_lb_controller_tag_version: v0.16.0
metal_lb_interfaces: []
metal_lb_ip_range: 192.168.30.80-192.168.30.90
+8
View File
@@ -141,6 +141,14 @@ argument_specs:
description: MetalLB ip range for load balancer
default: 192.168.30.80-192.168.30.90
metal_lb_interfaces:
description: >-
List of network interfaces on which MetalLB should announce the
load balancer IPs in layer2 mode. When empty (default), MetalLB
announces on all interfaces.
type: list
default: []
metal_lb_controller_tag_version:
description: Image tag for MetalLB
default: v0.16.0
+1
View File
@@ -178,6 +178,7 @@
--helm-set hubble.enabled={{ "true" if cilium_hubble else "false" }}
--helm-set hubble.relay.enabled={{ "true" if cilium_hubble else "false" }}
--helm-set hubble.ui.enabled={{ "true" if cilium_hubble else "false" }}
--helm-set envoy.enabled={{ "true" if cilium_envoy else "false" }}
{% if kube_proxy_replacement is not false %}
--helm-set loadBalancer.algorithm={{ bpf_lb_algorithm }}
--helm-set loadBalancer.mode={{ bpf_lb_mode }}
+1 -1
View File
@@ -40,7 +40,7 @@
- name: Test metallb-system namespace
ansible.builtin.command: >-
{{ k3s_kubectl_binary | default('k3s kubectl') }} -n metallb-system
{{ k3s_kubectl_binary | default('k3s kubectl') }} get namespace metallb-system
changed_when: false
with_items: "{{ groups[group_name_master | default('master')] }}"
run_once: true
@@ -21,6 +21,11 @@ kind: L2Advertisement
metadata:
name: default
namespace: metallb-system
{% if metal_lb_interfaces | default([]) | length > 0 %}
spec:
interfaces:{% for iface in metal_lb_interfaces %}
- {{ iface }}{% endfor %}
{% endif %}
{% endif %}
{% if metal_lb_mode == "bgp" %}
---
+2
View File
@@ -1,4 +1,6 @@
---
disable_swap: true
secure_path:
RedHat: /sbin:/bin:/usr/sbin:/usr/bin:/usr/local/bin
Suse: /usr/sbin:/usr/bin:/sbin:/bin:/usr/local/bin
+22
View File
@@ -4,6 +4,28 @@
name: "{{ system_timezone }}"
when: (system_timezone is defined) and (system_timezone != "Your/Timezone")
# k3s recommends swap be disabled on all nodes. Disabling swap is all-or-nothing
# across the cluster: leaving it enabled on some nodes but not others creates
# uneven scheduling/latency behavior. This block turns swap off and comments out
# the swap entries in /etc/fstab so it stays off across reboots. It is idempotent
# and a no-op when swap is already disabled or swapoff is unavailable.
- name: Disable swap on all cluster nodes
when: disable_swap
block:
- name: Turn off swap now
ansible.builtin.command: swapoff -a
register: swapoff_result
changed_when: false
failed_when: false
- name: Comment out swap entries in fstab
ansible.builtin.replace:
path: /etc/fstab
regexp: '^([^#][^\n]*\s+swap\s+)'
replace: '# \\1'
register: fstab_swap
- name: Set SELinux to disabled state
ansible.posix.selinux:
state: disabled
+17
View File
@@ -8,6 +8,23 @@
msg: >
"Ansible is out of date. See here for more info: https://docs.technotim.com/posts/ansible-automation/"
- name: Verify all cluster nodes have unique hostnames
ansible.builtin.assert:
that: (cluster_hostnames | unique | length) == (cluster_hostnames | length)
msg: >-
k3s nodes must have unique hostnames. Found a duplicate in:
{{ cluster_hostnames | unique }}. Each node registers in the cluster
keyed by its hostname, so matching hostnames prevent nodes from joining.
vars:
cluster_hostnames: >-
{{
groups['k3s_cluster']
| map('extract', hostvars, 'ansible_hostname')
| list
}}
run_once: true
when: "'k3s_cluster' in groups"
- name: Prepare Proxmox cluster
hosts: proxmox
gather_facts: true