Failover
Chronicle FIX failover provides fast and efficient recovery when a connection to the primary acceptor fails. To support failover, backup (secondary) acceptors receive replicas of all outgoing FIX messages from the primary acceptor to the initiator engine. In the event of a primary acceptor failure, the secondary acceptor can continue to exchange messages with the initiator. Chronicle Queue replication is used to create these replicas, which contain all the state information used by the Chronicle FIX session, including message sequence numbers.
Additionally, after a connection failure, the initiator must establish a new connection with a secondary acceptor. Chronicle FIX supports automatic reconnection of initiator and acceptor through the ConnectionStrategy property. Figure 1 shows an overview of the Chronicle failover procedure.

1. Failover Enabler Mechanisms
Failover Enabler Mechanisms
To enable failover, two mechanisms are required: replication of FIX messages from the primary acceptor to secondary acceptors, and a reconnection strategy. These mechanisms are detailed in the following sections. See runnable examples of failover in the Chronicle FIX Demo Failover Feature example.
1.1. Replication of Message Store Queue
To assume the primary role after a connection failure between the initiator and primary acceptor, a secondary acceptor must be aware of the state of communication before the failure. This is achieved by replicating the message store queue of the primary acceptor to the secondary acceptor’s message store queue. By default, FIX messages are stored in a Chronicle Queue at the location defined by the FileStorePath configuration parameter (see Logging FIX Messages for more information). Therefore, replicating the message store queue from the primary acceptor to secondary acceptors is sufficient. Refer to the Chronicle Queue replication documentation for more details.
1.1.1. Replication Acknowledgment Strategies
A FIX session can be configured to use a replication acknowledgment strategy to balance consistency across the replication cluster against message latency. The role of the replication acknowledgment strategy is to instruct Chronicle FIX when to wait before sending a message to a FIX counterparty. For example, Chronicle FIX may wait for acknowledgment that replication has reached another host.
The predefined strategies are NEVER_TIMES_OUT and WAIT_FOREVER_FOR_REPLICA_ACK, described below.
ReplicationAcknowledgementStrategy.NEVER_TIMES_OUT
This strategy does not wait for replication acknowledgment and does not prevent the FIX engine from sending messages. It provides optimal performance but carries the risk of message loss if the primary engine abruptly terminates before the secondary engine or counterparty has received the message. NEVER_TIMES_OUT is the default strategy.
ReplicationAcknowledgementStrategy.WAIT_FOREVER_FOR_REPLICA_ACK
This strategy ensures the FIX engine only sends a message to the counterparty once it has been acknowledged as replicated, guaranteeing zero message loss even if the primary engine terminates abruptly.
The ReplicationAcknowledgementStrategy determines the trade-off between consistency and availability, with the two predefined strategies representing the extremes. NEVER_TIMES_OUT prioritizes availability, ensuring the counterparty always receives a response, but sacrifices consistency. WAIT_FOREVER_FOR_REPLICA_ACK guarantees consistency between the primary and secondary message store queues at the cost of availability, as no messages are sent to the counterparty if replication between replicas stops.
1.2. Connection Strategy
ConnectionStrategy and socketConnectHostPort are configuration parameters that control how an initiator connects to acceptors. For failover, these parameters determine the secondary acceptor following a connection failure and enable automatic reconnection. Both parameters must be set for sessions with initiator ConnectionType, even if a failover mechanism is not implemented. Read more about these parameters and their configuration in the Connection Strategy documentation.
2. Failover Considerations
2.1. Network Address Translation
We recommend that your organisation uses NAT (Network Address Translation) so that the external address provided to clients differs from your internal address. This approach offers additional security and ensures clients do not need to be aware of your primary and secondary hosts. Clients should always connect to the same address, regardless of whether failover has occurred. Therefore, re-routing NATs is a key component of supporting external failover.
2.2. Manual Failover Between Acceptors
For scenarios requiring manual control over the switchover of acceptors, follow these steps:
Stop the primary acceptor FIX engine. The secondary acceptor FIX engine should not be started yet.
Stop replication (both source and sink replicators) and wait for confirmation that both have stopped.
Swap the source and sink roles of the queues and restart replication.
Start the secondary acceptor FIX engine.
It is possible to have both primary and secondary FIX engines started in step 1, but the secondary sessions should not start (set autoStart = false and do not call FixSessionHandler.start()). In step 4, explicitly start the secondary sessions.
2.3. Outage Detection and Leader Election
The primary node should be the one clients connect to, and they should only connect to a secondary node if the primary fails. If the primary fails, a load balancer can redirect traffic to the secondary, but most FIX connectors support round-robin connections across multiple hostnames or IP addresses.
While it is possible to run a fully self-governing cluster of Acceptors, it is typically preferable to retain a manual signal to coordinate failover to a new primary Acceptor. This approach is often chosen because:
Immediate failover may not always be appropriate; in some cases, restarting the failed primary may be less disruptive to the business or clients.
It avoids split-brain issues, where network problems prevent all nodes in a resilient cluster from communicating, potentially resulting in two nodes attempting to assume primary status.