--- title: "High Availability (HA) Cluster" description: "Documentation for High Availability (HA)" --- ## Table of Contents 1. [Overview & High Availability Architecture](#1-overview--high-availability-architecture) 2. [Business & Operational Significance](#2-business--operational-significance) 3. [🎯 User Roles & Key Capabilities](#3--user-roles--key-capabilities) 4. [Visual Interface & Layout](#4-visual-interface--layout) 5. [Field Reference & Node Parameters](#5-field-reference--node-parameters) 6. [Kamailio DMQ Replication Protocol Mechanics](#6-kamailio-dmq-replication-protocol-mechanics) 7. [Failover Strategies: VRRP VIP vs Anycast / DNS SRV](#7-failover-strategies-vrrp-vip-vs-anycast--dns-srv) 8. [Troubleshooting & Cluster Verification](#8-troubleshooting--cluster-verification) 9. [Model Context Protocol (MCP) AI Integration](#9-model-context-protocol-mcp-ai-integration) 10. [Glossary](#10-glossary) --- ## 1. Overview & High Availability Architecture In **Ring2All SBC**, the **High Availability (HA) Cluster** module manages multi-node redundancy, real-time memory synchronization, and automatic failover. Utilizing Kamailio's native **Distributed Message Queue (DMQ)** protocol over a dedicated internal network port (default UDP/5062), the SBC synchronizes active SIP dialog states, user registration bindings (`usrloc`), and security hash tables (`htable`) continuously across all nodes in the cluster. ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Shared Virtual IP / DNS SRV Weight β”‚ β”‚ 192.168.10.30 β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ sbc-node-01 β”‚ β”‚ sbc-node-02 β”‚ β”‚ (Master / Priority 100)β”‚ β”‚ (Standby / Priority 90) β”‚ β”‚ 192.168.10.32 β”‚ β”‚ 192.168.10.33 β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ SIP KDMQ Sync (UDP 5062) β”‚ │◄══════════════════════════════►│ β”‚ β€’ Active Dialog States β”‚ β”‚ β€’ HTables & IP Bans β”‚ β”‚ β€’ User Registrations β”‚ ``` If the primary node experiences hardware failure, network isolation, or kernel panic, the secondary node immediately takes over the virtual IP and assumes ownership of all active SIP calls without terminating conversational audio or dropping carrier interconnects. --- ## 2. Business & Operational Significance * **Five-Nines Availability (99.999%)**: Delivers carrier-grade telecom reliability by eliminating single points of failure across the perimeter signaling plane. * **In-Flight Call Survivability**: By replicating dialog state machines in real-time, secondary nodes can process `BYE`, `re-INVITE`, and `INFO` packets for calls that were established by the primary node before failover. * **Synchronized Perimeter Security**: Dynamic security events (such as Pike anti-flood bans or APIBAN firewall drops) discovered on one node are broadcast instantly to all cluster nodes, protecting the entire edge simultaneously. * **Hitless Maintenance & Upgrades**: Allows operators to gracefully drain traffic from a node, perform kernel or package updates, and rejoin the cluster without customer-facing downtime. --- ## 3. 🎯 User Roles & Key Capabilities | Role | Primary Use Case | Key Capabilities | | :--- | :--- | :--- | | **SBC High Availability Engineer** | Topology & Redundancy Design | Configure cluster node priorities, manage DMQ interconnect listeners, and define heartbeat timeouts. | | **Infrastructure Architect** | Network Failover Coordination | Align SBC cluster nodes with underlying VRRP (`keepalived`) virtual IPs or BGP Anycast route health injectors. | | **Telecom Systems Administrator** | Node Provisioning & Maintenance | Add new secondary nodes, monitor node health badges, and initiate graceful drain procedures for software updates. | | **NOC Tier 3 Operations** | Real-Time Cluster Triage | Investigate node synchronization status, execute DMQ ping diagnostics, and audit inter-node latency. | | **AI Infrastructure Copilot / NOC Agent** | Topology & Sync Telemetry | Audit multi-node DMQ replication health, probe inter-node socket latencies, and trigger state resynchronization via MCP. | --- ## 4. Visual Interface & Layout The HA Cluster interface consists of a centralized cluster management view displaying all registered nodes, their replication health, and priorities, along with a dedicated modal for node registration and tuning. ### 4.1 Cluster Nodes List View Displays all nodes in the mesh topology, their DMQ and SIP listening sockets, replication roles, priority weights, and real-time connectivity status. ![High Availability Cluster Nodes List View](/screenshots/sbc/settings/technology/ha-cluster/ha-cluster-list.png) ### 4.2 Node Configuration Form Modal used to register or adjust node network identities, ports, priority weights, and health check intervals. ![HA Cluster Node Configuration Form](/screenshots/sbc/settings/technology/ha-cluster/ha-node-form.png) --- ## 5. Field Reference & Node Parameters | Parameter | Data Type | Default | Description | | :--- | :--- | :--- | :--- | | **Node Name** | String | `sbc-node-01` | Unique hostname or identifier for the cluster instance. | | **IP Address** | IPv4 / IPv6 | `192.168.10.32` | Dedicated internal network address used for inter-node DMQ traffic and SIP synchronization. | | **DMQ Port** | Integer | `5062` | UDP port allocated strictly for Kamailio `dmq` and `dmq_usrloc` mesh message exchange. | | **SIP Port** | Integer | `5060` | Public or internal SIP signaling port on which this node processes external traffic. | | **Role** | Select | `Master` | Operational state: `Master` (primary traffic processor) or `Standby` (hot standby replica). | | **Priority** | Integer | `100` | Integer weight (1-100) determining master election order. The highest responsive priority node assumes the master state. | | **Heartbeat Interval** | Integer (ms) | `1000` | Frequency in milliseconds at which nodes exchange `KDMQ` ping probes to verify peer liveness. | | **Health Status** | Badge | `Healthy` | Real-time state evaluated by DMQ peer responses (`Healthy` Green, `Degraded` Yellow, `Offline` Red). | --- ## 6. Kamailio DMQ Replication Protocol Mechanics Kamailio’s DMQ module implements a peer-to-peer distributed message bus inside the SIP engine. Rather than relying on slow SQL replication, DMQ transmits binary SIP requests using the `KDMQ` method: ```text # Sample Kamailio DMQ module configuration loadmodule "dmq.so" loadmodule "dmq_usrloc.so" modparam("dmq", "server_address", "sip:192.168.10.32:5062") modparam("dmq", "notification_address", "sip:192.168.10.33:5062") modparam("dmq", "multi_notify", 1) modparam("dmq", "ping_interval", 1) ``` When a new call is established on `sbc-node-01`, the `dialog` module serializes the session state (Call-ID, From-tag, To-tag, Caller/Callee Contact, RTPEngine call ID) and delivers it across the DMQ socket. `sbc-node-02` immediately inserts this session into its shared memory dialog table. --- ## 7. Failover Strategies: VRRP VIP vs Anycast / DNS SRV Ring2All SBC supports two primary deployment topologies for carrier interconnects: 1. **Virtual IP (VRRP / Keepalived)**: * Both nodes share a floating IP (e.g., `192.168.10.30`). * The Master node binds the VIP. If the Master stops responding to heartbeats, Keepalived transfers the VIP to the Standby node in `< 800ms`. * Carriers send all SIP traffic to the single VIP without needing multi-IP failover logic. 2. **DNS SRV / RFC 3263 Multi-Record**: * Carriers query DNS SRV records for `_sip._udp.sbc.ring2all.com`. * Node 01 has weight 100, Node 02 has weight 10. * If Node 01 fails ICMP/SIP probes, carrier trunks automatically divert new `INVITE` attempts to Node 02. --- ## 8. Troubleshooting & Cluster Verification ### Inspecting DMQ Peers via RPC Console Execute the following commands in the **RPC Console**: ```bash dmq.list_nodes ``` Expected output: ```json { "jsonrpc": "2.0", "result": [ { "host": "192.168.10.33", "port": 5062, "status": "Active", "last_ping": "1s ago", "latency_ms": 0.45 } ], "id": 1 } ``` ### Forcing a Manual DMQ Ping Probe ```bash dmq.ping ``` --- ## 9. Model Context Protocol (MCP) AI Integration The High Availability subsystem exposes dedicated Model Context Protocol (MCP) tools enabling intelligent monitoring agents to inspect multi-node cluster health, verify inter-node socket reachability, and force state synchronization. ### Available MCP Tools | Tool Name | Operation Type | Risk Level | Description | | :--- | :--- | :--- | :--- | | `get_ha_cluster_status` | Status Query | `read` | Retrieve cluster topology, Virtual IP assignment, Kamailio DMQ engine state, and replicated subsystems. | | `list_cluster_nodes` | Status Query | `read` | List all peer SBC nodes in the DMQ mesh with real-time SIP and DMQ socket reachability and latency metrics. | | `sync_cluster_state` | Operational Sync | `operational` | Force an immediate Kamailio DMQ state synchronization broadcast across all registered cluster members. | ### Tool Schemas & Payloads #### 1. `get_ha_cluster_status` ##### Input Schema ```json { "type": "object", "properties": {} } ``` ##### Output Payload Example ```json { "success": true, "data": { "clusterMode": "Active-Active Mesh (DMQ)", "totalNodesConfigured": 2, "virtualIp": "192.168.10.30", "dmqSyncPort": 5062, "replicatedSubsystems": [ "SIP Dialogs (dialog)", "User Location (usrloc)", "Memory Tables (htable)", "RTPEngine Media Sessions" ], "dmqNodesRaw": "Nodes: 2, Active: 2, Synced: Yes" } } ``` #### 2. `list_cluster_nodes` ##### Input Schema ```json { "type": "object", "properties": {} } ``` ##### Output Payload Example ```json { "success": true, "data": { "totalNodes": 2, "nodes": [ { "id": 1, "name": "sbc-node-01", "role": "Master", "ip": "192.168.10.32", "sipPort": 5060, "dmqPort": 5062, "isSelf": true, "status": "ONLINE", "sipLatency": "0.1ms", "dmqLatency": "0.1ms" }, { "id": 2, "name": "sbc-node-02", "role": "Standby", "ip": "192.168.10.33", "sipPort": 5060, "dmqPort": 5062, "isSelf": false, "status": "ONLINE", "sipLatency": "0.45ms", "dmqLatency": "0.52ms" } ] } } ``` #### 3. `sync_cluster_state` ##### Input Schema ```json { "type": "object", "properties": {} } ``` ##### Output Payload Example ```json { "success": true, "data": { "message": "Cluster state replication synchronized across Kamailio SBC nodes!", "details": "DMQ sync OK" } } ``` ### Natural Language AI Prompts #### English Examples * *"Check the high availability cluster status and verify if all peer nodes are synchronized."* * *"List all SBC cluster nodes with their SIP and DMQ socket latencies."* * *"Broadcast an immediate DMQ cluster sync to ensure standby nodes have active dialog tables."* #### Spanish Examples (EspaΓ±ol) * *"Verifica el estado del cluster de alta disponibilidad y comprueba si todos los nodos pares estΓ‘n sincronizados."* * *"Lista todos los nodos del cluster SBC con las latencias de sus sockets SIP y DMQ."* * *"Emite una sincronizaciΓ³n DMQ inmediata del cluster para asegurar que los nodos en standby tengan las tablas de diΓ‘logos activas."* ### Enterprise Safeguards & Access Governance 1. **Isolated Health Probing**: Socket latency checks utilize short 2000ms timeouts to avoid hanging diagnostic agent threads. 2. **Master Election Protection**: MCP tools do not unilaterally force master reassignment; elections follow deterministic VRRP priorities. 3. **Non-Disruptive Sync**: Calling `sync_cluster_state` uses native Kamailio asynchronous DMQ multicasting without dropping active SIP calls. --- ## 10. Glossary * **DMQ (Distributed Message Queue)**: An internal Kamailio protocol module that facilitates sub-millisecond in-memory state replication across independent SBC nodes. * **KDMQ Method**: A custom SIP request method used exclusively for inter-node communication and state synchronization. * **VRRP (Virtual Router Redundancy Protocol)**: A computer networking protocol that provides automatic assignment of available IP routers to participating hosts. * **Dialog Replication**: Copying active session metadata between SBC nodes so that call termination and in-dialog requests can be fulfilled by any node in the cluster.