HA 内存高触发切换
HA 内存高触发切换
在 HA 部署中,如果主设备存在持续的内存泄漏或异常内存占用,可以启用基于内存的 HA 故障转移,使业务迁移到内存占用较低的成员。HA 切换会影响流量处理,应先确认内存问题确实需要通过切换缓解,并在维护窗口验证。
主备选举顺序
开启 memory-based-failover 后,HA 主备选举依次参考以下条件:
- 有效监控接口数量。
mem_failoverflag 状态。- HA 有效运行时间。
- HA 优先级。
- 序列号。
工作机制
- 内存使用率超过
memory-failover-threshold后,FortiGate 按memory-failover-sample-rate进行采样。在整个memory-failover-monitor-period内每次采样都超过阈值时,才触发 HA 切换,业务迁移到备机,mem_failover标志从0变为1,并生成 HA 事件日志。 - 内存回落且经过
memory-failover-flip-timeout配置的等待时间后,mem_failover标志从1变为0,HA 可能再次进行选举。同一成员刚发生过基于内存的切换时,在该等待时间内不会再次触发同类切换,等待时间结束后仍需重新满足memory-failover-monitor-period才可能再次切换。 - 第一次内存触发切换后,可以使用自动化执行
diagnose sys ha reset-uptime,重置 HA uptime,防止因mem_failover标志状态恢复而由 uptime 导致的反向选举。
HA 参数配置
memory-failover-threshold 应低于 Conserve Mode 的 red 阈值,在设备进入内存保护模式前先触发 HA 切换。例如设备的 red 阈值为 88% 时,可将内存故障转移阈值设置为 85%:
config system ha
set memory-based-failover enable
set memory-failover-threshold 85
set memory-failover-monitor-period 10
set memory-failover-sample-rate 1
set memory-failover-flip-timeout 6
endmemory-based-failover:启用基于内存的 HA 故障转移。memory-failover-threshold:触发内存故障转移的内存使用率,范围为0~95。示例值为85,应低于red阈值。memory-failover-monitor-period:触发故障转移前,内存使用率持续超过阈值的时间,示例为10秒。memory-failover-sample-rate:内存采样间隔,示例为1秒。在整个监控周期内,每次采样都超过阈值时才触发故障转移。memory-failover-flip-timeout:连续两次基于内存故障转移之间的等待时间,示例为6分钟。
重要
不要把 memory-failover-threshold 设置为 red 阈值或更高,否则 HA 切换可能与 Conserve Mode 同时发生,失去提前迁移业务的缓冲区。
自动化配置
创建触发器。内存故障转移产生的事件日志 ID 为
37904。触发器还需要精确匹配activity字段,避免把清除mem_failover标志的日志也作为触发条件。config system automation-trigger edit "Memory based failover flag is set" set event-type event-log set logid 37904 config fields edit 1 set name "activity" set value "Memory based failover flag is set" next end next end创建动作。内存恢复后,原主设备成为备机,双方的
mem_failoverflag 可能都恢复为0。如果原主设备的 HA uptime 更长,下一轮选举可能重新抢占主设备。通过diagnose sys ha reset-uptime重置目标备机的 HA uptime,可以降低其因 uptime 优势重新抢主的可能,让内存占用较低的新主继续承载业务。config system automation-action edit "Reset HA uptime" set action-type cli-script set script "diagnose sys ha reset-uptime" set accprofile "super_admin_readonly" next end创建 Stitch。
config system automation-stitch edit "Memory based failover" set trigger "Memory based failover flag is set" config actions edit 1 set action "Reset HA uptime" set required enable next end next end
状态与日志验证
执行
diagnose hardware sysinfo conserve,记录当前 Conserve Mode 状态和三个内存阈值。# diagnose hardware sysinfo conserve memory conserve mode: off total RAM: 24140 MB memory used: 5381 MB 22% of total RAM memory freeable: 312 MB 1% of total RAM memory used + freeable threshold extreme: 22933 MB 95% of total RAM memory used threshold red: 21243 MB 88% of total RAM memory used threshold green: 19795 MB 82% of total RAM对照
memory-failover-threshold与memory used threshold red,确认前者低于red,使 HA 能在进入 Conserve Mode 前触发切换。查看 HA 成员状态、选举原因和 HA 事件日志,确认
mem_failover的设置和清除状态,并判断是否触发了新一轮 HA 选举。执行以下命令查看每个成员的
mem_failover、uptime/reset_cnt、ha_prio、监控接口状态和pingsvr_flip_timeout。其中,mem_failover=1表示该成员发生过基于内存的 HA 切换,另一成员通常为mem_failover=0。# diagnose sys ha dump-by group vcluster_0: start_time=1645078182(2022-02-17 14:09:42), state/o/chg_time=2(work)/3(standby)/1645079708(2022-02-17 14:35:08) pingsvr_flip_timeout/expire=3600s/3452s mondev: port1(prio=50,is_aggr=0,status=1) port2(prio=50,is_aggr=0,status=1) 'FG5H1E5819904036': ha_prio/o=0/0, link_failure=0, pingsvr_failure=0, flag=0x00000001, mem_failover=0, uptime/reset_cnt=15/0 'FG5H1E5819904211': ha_prio/o=1/1, link_failure=0, pingsvr_failure=0, flag=0x00000040, mem_failover=1, uptime/reset_cnt=0/0执行以下命令查看当前主设备的选举原因。输出中的
is selected as the primary because there is high memory usage on peer member表示因对端成员内存使用率较高而成为主设备,is selected as the primary because its uptime is larger than peer member表示因 HA uptime 更长而成为主设备。# get system ha status HA Health Status: OK Model: FortiGate-501E Mode: HA A-P Group: 0 Debug: 0 Cluster Uptime: 0 days 1:19:20 Cluster state change time: 2022-02-17 15:27:11 Primary selected using: <2022/02/17 15:27:11> FG5H1E5819904036 is selected as the primary because its uptime is larger than peer member FG5H1E5819904211. <2022/02/17 14:59:48> FG5H1E5819904211 is selected as the primary because there is high memory usage on peer member FG5H1E5819904036. <2022/02/17 14:49:08> FG5H1E5819904036 is selected as the primary because there is high memory usage on peer member FG5H1E5819904211. <2022/02/17 14:40:38> FG5H1E5819904211 is selected as the primary because there is high memory usage on peer member FG5H1E5819904036.查看 HA 事件日志,确认
mem_failover的设置和清除状态。内存回落后,原主设备成为备机,双方的mem_failoverflag 都可能恢复为0。如果需要避免备机因 HA uptime 优势重新抢占主设备,可以在目标备机上执行diagnose sys ha reset-uptime。上面的自动化用于在内存故障转移事件发生后自动执行该命令。date=2026-06-04 time=14:00:07 eventtime=1780552807798851106 tz="+0800" logid="0108037904" type="event" subtype="ha" level="notice" vd="root" logdesc="Device set as HA primary" msg="HA activity report" activity="Memory based failover flag is set" date=2026-06-04 time=14:37:10 eventtime=1780555030023035648 tz="+0800" logid="0108037904" type="event" subtype="ha" level="notice" vd="root" logdesc="Device set as HA primary" msg="HA activity report" activity="Memory based failover flag is cleared"