- Add shared download_retries/download_delay defaults in k3s_server and k3s_server_post roles
- Retry calico CRD and Tigera operator manifest downloads
- Retry Cilium CLI download and cilium install/upgrade command
- Retry kube-vip cloud provider and MetalLB manifest downloads
- The CI runner's resolver intermittently times out on GitHub-hosted domains
- Correct the Deploy metallb manifest/pool when condition so MetalLB is
installed whenever kube-vip does not own the VIP range and Cilium BGP
is disabled
- The previous guard (cilium_bgp is not defined or cilium_iface is not
defined) skipped MetalLB whenever cilium_iface was set, breaking the
cilium + MetalLB scenario
- Use cilium_bgp | default(false) | bool to stay safe when Cilium vars are
not in scope (#644) while still deploying MetalLB for non-BGP Cilium
- Retry the converge-side MetalLB resource wait so a transient kube API
ServiceUnavailable does not abort the converge play
- Add a regression test that evaluates both when conditions across flannel,
calico, non-BGP cilium, BGP cilium, and kube-vip scenarios
ci: skip CI for Dependabot pull requests
- Add an actor guard to the CI workflow jobs so automatic Dependabot PRs
do not consume the shared self-hosted runner
- Dependabot CI runs need maintainer approval instead of auto-running
* fix(flannel): default the interface to each host's default IPv4 interface
- Replace the hardcoded flannel_iface: eth0 with a per-host default derived
from ansible_facts.default_ipv4.interface
- KVM/cloud hosts that are not named eth0 (e.g. enp1s0, ens3) now resolve the
interface automatically instead of failing the k3s_node_ip lookup
- Update the commented calico_iface / cilium_iface examples to match
- Co-authored-by: Fritz Dunkel <677609+FinalDoom@users.noreply.github.com>
- Fixes#621
* test(flannel): add regression test for per-host interface default
- Add a focused test that renders the sample inventory's flannel_iface
expression against fake ansible facts
- Assert a non-eth0 host (enp1s0, ens3) resolves its own interface and that
the expression defaults from ansible_facts.default_ipv4.interface
- Wire it as a local pre-commit hook (default-interface-test)
* test(molecule): assert nodes register a non-loopback InternalIP
- Add a flannel-scenario verify assertion that every node reports an
InternalIP derived from its configured flannel_iface
- Guards against k3s binding to 127.0.0.1 instead of the cluster interface
- Complements the pre-commit default-interface test with an end-to-end check
- On RHEL 10 / Rocky Linux 10 the default cloud image does not ship the
kernel-modules-extra that contains br_netfilter, so modprobe br_netfilter
fails during prereq on RedHat family
- Detect whether the module exists and install kernel-modules-extra when absent
- No reboot is required: the module is installed for the currently running kernel
and becomes available to modprobe immediately
* fix(metallb): guard cilium_bgp variable before evaluating
- The 'Deploy metallb manifest' and 'Deploy metallb pool' conditionals evaluate
'not cilium_bgp' directly, which raises an undefined-variable error when the
k3s_server role runs without cilium_bgp in scope and no Cilium variables are set
- Guard with 'cilium_bgp is not defined' so the condition resolves cleanly when
Cilium BGP is not configured
- Fixes#644
* docs(apiserver): clarify apiserver_endpoint must be a free routable IP
- Note that apiserver_endpoint must be an unassigned, routable IP on the
network and that it is exposed by kube-vip / MetalLB
- Fixes#678
* fix(molecule): wait for MetalLB resources before asserting images
- The MetalLB image-tag assertion crashed with 'list object has no element 0'
when the controller Deployment was not yet observable at verify time
- Retry the MetalLB controller/speaker lookup until the resources appear
- Fail with a clear message if MetalLB is genuinely absent
- Add canonical agent and contributor documentation.
- Modernize issue and pull request templates.
- Skip CI for documentation and issue-template-only changes.
- systemd treats % as a specifier in ExecStart, so token values containing %
fail with 'failed to resolve unit specifiers invalid slot'
- escape % to %% in the agent unit file
- escape % to %% in server_init_args so the multi-master bootstrap path
(systemd-run k3s-init) handles % tokens too
Co-authored-by: finaldoom <677609+FinalDoom@users.noreply.github.com>
* fix(calico): support split CRDs for current releases
- Download the v1_crd_projectcalico_org.yaml bundle before the operator
- Apply both files with server-side apply and force-conflicts per the
upstream upgrade procedure
- Wait for the operator Deployment and for the managed CRDs to be
Established after the operator starts
- Replace the create/rescue/replace flow with an idempotent apply that
no longer conceals partial failures
- Verify TigeraStatus for calico and apiserver is Available, not just
that Pods exist
* feat(dependencies): upgrade supported cluster components
- Bump K3s to v1.36.2+k3s1, Calico to v3.32.1, Cilium to v1.20.0,
kube-vip to v1.2.2, kube-vip cloud provider to v0.0.12, and MetalLB to
v0.16.0 across sample inventory, role defaults, and argument specs
- Pin the Cilium CLI with a new cilium_cli_tag (v0.19.7) instead of the
floating stable.txt lookup
- Replace the CiliumBGPPeeringPolicy v2alpha1 BGP template with the
v2 CiliumBGPClusterConfig, CiliumBGPPeerConfig, CiliumBGPAdvertisement,
and CiliumLoadBalancerIPPool resource set
- Move Cilium load balancer Helm keys from bpf.loadBalancer to the valid
top-level loadBalancer path
- Add preflight schema validation and remove the deprecated policy after
the v2 objects are accepted
- Wait for cilium status after installation
- Pin kube-vip RBAC in a repository template instead of fetching a
mutable URL, and include EndpointSlice permissions
- Fix the kube-vip bgppeers format to address:ASN comma-separated peers
- Fail clearly when the MetalLB speaker tag replacement does not apply
- Drop the obsolete MetalLB webhook service name version branch
* test(molecule): verify upgraded cluster components
- Assert every node reports the expected K3s kubelet version
- Verify the active CNI (Flannel / Calico / Cilium) is Ready and runs
the expected image tag, including Calico TigeraStatus Available
- Verify the active load balancer (MetalLB / kube-vip) runs the expected
image tags and that MetalLB is absent when kube-vip is active
- Assert no Flannel DaemonSet remains when Calico or Cilium is enabled
- Assert the example LoadBalancer address falls inside the configured
pool range
- Add a manifest-only Cilium BGP regression test that renders the v2
template with zero, one, and multiple neighbors and rejects any v2alpha1
or CiliumBGPPeeringPolicy output
* fix(dependencies): correct dependency version pins
- Set the sample kube-vip image to v1.2.2 and repair the damaged comment
- Pin the kube-vip cloud provider default to v0.0.12 in the task URL
- Set the MetalLB controller argument-spec default to v0.16.0
- Restore the MetalLB available timeout default to 240s
* docs(dependencies): document current cluster versions
- Update kube-vip, kube-vip cloud provider, and MetalLB defaults
- Add cilium_tag and cilium_cli_tag rows
- Explain that MetalLB v0.16.0 is the application image target even though
a newer chart-only tag (metallb-chart-0.16.1) exists
- Add an existing-cluster upgrade warning covering the K3s etcd 3.5.26
bridge and one-minor-at-a-time rule, consecutive Cilium minor upgrades,
Calico v3 resource UID handling, and MetalLB app vs chart tags
* fix(dependencies): address PR review findings
- Read the MetalLB speaker tag check from the managed host with slurp
instead of a controller-side file lookup, and match the full image
reference
- Restore the tigera-operator namespace on the Calico operator Deployment
wait while keeping the managed CRD waits cluster-scoped
- Make Molecule verify inputs durable and scenario-specific via a
per-scenario verify-vars.yml, driven by explicit verify_cni/verify_lb
values instead of non-persisted converge facts
- Rename the kube-vip multi-peer BGP env var from bgppeers to bgp_peers
and vip_cidr to vip_subnet so v1.2.2 actually reads them
- Map the legacy Cilium routed mode to tunnel and stop passing the alias
directly to the chart
- Use return-code based failed_when on apply and preflight commands so
non-error failures are no longer treated as success
- Clarify the sequential K3s upgrade path and backups in the README
- Add kube-vip and MetalLB regression tests and a Cilium mode mapping unit
* fix(dependencies): resolve re-review findings
- correct the Calico TigeraStatus resource kind\n- document tunnel as the supported Cilium routing mode\n- validate load balancer addresses across range and CIDR pools
* fix(molecule): verify embedded flannel instead of a flannel DaemonSet
- K3s 1.36 runs flannel embedded in the k3s agent rather than as a
kube-flannel-ds DaemonSet, so the flannel verifier queried a workload
that no longer exists and failed the verify step
- For the flannel scenarios, assert every node is Ready and that neither
the Calico nor the Cilium namespace exists
- Drop the now-invalid kube-flannel-ds DaemonSet assertion
* fix(molecule): wait for the LoadBalancer address before asserting reachability
- The nginx LoadBalancer service had no ingress address when the
reachability assertion ran, so status.loadBalancer.ingress[0].ip was
undefined and the ipwrap filter failed during verify
- Poll the service until MetalLB or kube-vip assigns an external IP
- Record the assigned address once and reuse it for the reachability probe
and the pool membership checks
* fix(ci): harden calico apiserver wait and extend molecule job timeout
- Bump calico system resources wait retries 30->60 and delay 7->10 so the
slow-to-reconcile calico-apiserver deployment has enough time under nested-virt
- Raise the molecule step timeout-minutes from 90 to 150 to accommodate
contended 5-node scenarios (cilium, kube-vip) that were hitting the 90-min cap
* fix(calico): treat optional API server as best-effort on converge
- The Calico API server (calico-apiserver) is an optional add-on for managing
Calico policy through the projectcalico.org/v3 Kubernetes API; it is not
required for Calico CNI data plane operation
- With Calico v3.32.1 on K3s 1.36 the tigera-operator never provisions the
calico-apiserver namespace, causing the converge wait to fail deterministically
- Keep the strict wait for core Calico components (typha, kube-controllers,
calico-node, csi-node-driver) and make the API server wait tolerate failure
- Restrict the TigeraStatus Available check to the calico status, matching the
upstream v3.32.1 K3s quickstart which validates without the API server
- Reuse bounded Vagrant startup batches when machine state already exists.
- Avoid falling back to a concurrent five-guest startup during retries.
- Keep instance reconciliation after all batches complete.
- Avoid concurrent SSH startup bursts against the disposable guests.
- Ensure every guest gathers interface facts before network setup.
- Prevent unreachable hosts from falling back to the default eth0 interface.
- Use eth1 for Ubuntu and Debian Bento guests when present.
- Keep enp0s8 for Rocky Linux guests and other images.
- Apply the selection consistently across all Molecule scenarios.
- Replace legacy eth1 overrides with the enp0s8 private NIC used by Bento guests.\n- Keep flannel, Calico, Cilium, and kube-vip scenarios on the Vagrant private network.
- Use published Bento boxes for Ubuntu 26.04, Debian 13, and Rocky Linux 10.1.\n- Pin exact VirtualBox amd64 box versions in CI.\n- Remove the obsolete Ubuntu SSH workaround and update fixtures.
- Match the minimum CPU and memory used by Ubuntu Molecule platforms
- Restore Vagrant's normal key insertion after password authentication
- Cover the canonical prewarm configuration in the fixture
- Use the Molecule Ubuntu SSH settings when creating its linked-clone master
- Bound a prewarm SSH failure to ten minutes
- Verify the generated Vagrantfile in the master-preparation fixture
- delegate cgroups for transient K3s server units\n- verify inventory node registration without legacy role labels\n- wait for bootstrap CRDs before replacing the transient service
- restore host-only routing required by outside verification
- assign deterministic adapter MACs to the five default guests
- pin disposable full-mesh neighbor entries from live Ansible facts
- move the five-node cluster NICs to a named VirtualBox internal network
- preserve NAT connectivity for provisioning and downloads
- retain serialized creation and early neighbor identity validation
- serialize five-node Vagrant creation to avoid VirtualBox host-only races
- refresh and verify the primary neighbor mapping before convergence
- include interface MAC details in failure diagnostics
- pin kube-vip and cluster traffic to the private guest interface\n- disable disposable guest firewalls and verify API reachability before joins\n- keep control-plane orchestration on the primary and preserve failure diagnostics
- Preserve explicit per-host server initialization overrides\n- Build default join arguments from the delegated host variables\n- Keep initialization commands out of normal task output
- Retry transient k3s release downloads with bounded backoff.\n- Bound failure diagnostics and validate guest release connectivity.\n- Discover repository-owned Molecule state under the actual project root.
- Update checkout, setup-python, cache, upload-artifact, and pin enforcement actions\n- Keep every third-party action pinned to an immutable release commit\n- Replace the deprecated cache action reference that blocked CI
* docs: first modules' variable docs table
* docs: variables for k3s_server_post
* docs: lxc and prereq vars in README
* style: lint errors
* docs: argument_specs for proxmox_lxc
* docs: last variables found added to the README
With the kube_vip_bgp_peers it is possible to define
multiple BGP peer ASN & address pairs for kube-vip.
Sample:
```
kube_vip_bgp_peers:
- peer_address: 192.168.128.10
peer_asn: 64512
- peer_address: 192.168.128.11
peer_asn: 64512
- peer_address: 192.168.128.12
peer_asn: 64512
```
It is possible to merge further lists with kube_vip_bgp_peers__*
parameters.
Sample:
```
kube_vip_bgp_peers__extra:
- peer_address: 192.168.128.10
peer_asn: 64512
kube_vip_bgp_peers:
- peer_address: 192.168.128.11
peer_asn: 64512
- peer_address: 192.168.128.12
peer_asn: 64512
```
This will result in the following list of BGP peer ASN & address pairs:
```
- peer_address: 192.168.128.10
peer_asn: 64512
- peer_address: 192.168.128.11
peer_asn: 64512
- peer_address: 192.168.128.12
peer_asn: 64512
```
Signed-off-by: Christian Berendt <berendt@osism.tech>
Co-authored-by: Techno Tim <timothystewart6@gmail.com>