# Table of Contents

### Layer 2 Technologies

* [Ethernet Switching](/layer-2-technologies/ethernet-switching)
  * [L2 Switching Operations](/layer-2-technologies/ethernet-switching/l2-switch-operations)
  * [802.1d – STP](/layer-2-technologies/ethernet-switching/spanning-tree/802.1d-stp)
  * [802.1w – RSTP](/layer-2-technologies/ethernet-switching/spanning-tree/802.1w-rstp)
  * [802.1s – MSTP](/layer-2-technologies/ethernet-switching/spanning-tree/802.1s-mstp)
  * [VTP 101](/layer-2-technologies/ethernet-switching/vtp-101)
  * [Private VLANs](/layer-2-technologies/ethernet-switching/private-vlans)
  * [VLANs](/layer-2-technologies/ethernet-switching/vlans)
  * [EtherChannel 101](/layer-2-technologies/ethernet-switching/etherchannel-101)
* [Layer 2 WAN Protocols](/layer-2-technologies/layer-2-wan-protocols)
  * [HDLC](/layer-2-technologies/layer-2-wan-protocols/hdlc)
    * [HDLC 101](/layer-2-technologies/layer-2-wan-protocols/hdlc/hdlc-101)
  * [PPP](/layer-2-technologies/layer-2-wan-protocols/ppp)
    * [PPP 101](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-101)
    * [PPP Authentication - PAP](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-authentication-pap)
    * [PPP Authentication – CHAP](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-authentication-chap)
    * [PPP Authentication – EAP](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-authentication-eap)
    * [PPP Multilink](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-multilink)
    * [PPPoFR – PPP over Frame Relay](/layer-2-technologies/layer-2-wan-protocols/ppp/pppofr-ppp-over-frame-relay)
    * [PPPoE – PPP over Ethernet](/layer-2-technologies/layer-2-wan-protocols/ppp/pppoe-ppp-over-ethernet)
  * [Frame Relay](/layer-2-technologies/layer-2-wan-protocols/frame-relay)
    * [Frame Relay 101](/layer-2-technologies/layer-2-wan-protocols/frame-relay/frame-relay-101)
    * [Frame Relay 102](/layer-2-technologies/layer-2-wan-protocols/frame-relay/frame-relay-102)
    * [Frame Relay Encapsulations – IETF vs Cisco](/layer-2-technologies/layer-2-wan-protocols/frame-relay/frame-relay-encapsulations-ietf-vs-cisco)
    * [Multilink Frame Relay](/layer-2-technologies/layer-2-wan-protocols/frame-relay/multilink-frame-relay)
    * [Frame Relay Switching](/layer-2-technologies/layer-2-wan-protocols/frame-relay/frame-relay-switching)
    * [Routing over Frame Relay](/layer-2-technologies/layer-2-wan-protocols/frame-relay/routing-over-frame-relay)
  * [Bridging](/layer-2-technologies/layer-2-wan-protocols/bridging)
    * [Bridging on a router](/layer-2-technologies/layer-2-wan-protocols/bridging/bridging-on-a-router)
    * [MTU 101](/layer-2-technologies/layer-2-wan-protocols/bridging/mtu-101)

### IPv4

* [IPv4 Addressing](/ipv4/ipv4-addressing)
  * [Backup Interfaces](/ipv4/ipv4-addressing/backup-interfaces)
  * [FHRP 101](/ipv4/ipv4-addressing/fhrp-101)
  * [DHCP 101](/ipv4/ipv4-addressing/dhcp-101)
  * [DNS 101](/ipv4/ipv4-addressing/dns-101)
  * [ARP 101](/ipv4/ipv4-addressing/arp-101)
  * [IPv4 101](/ipv4/ipv4-addressing/ipv4-101)
  * [Tunnel Interfaces](/ipv4/ipv4-addressing/tunnel-interfaces)
    * [GRE Tunnels](/ipv4/ipv4-addressing/tunnel-interfaces/gre-tunnels)
  * [BFD – Bidirectional Forwarding Detection](/ipv4/ipv4-addressing/bfd-bidirectional-forwarding-detection)
* [IPv4 Routing](/ipv4/ipv4-routing)
  * [ODR](/ipv4/ipv4-routing/odr)
  * [Policy based Routing](/ipv4/ipv4-routing/policy-based-routing)
  * [How the routing table is built](/ipv4/ipv4-routing/how-the-routing-table-is-built)
    * [How CEF works](/ipv4/ipv4-routing/how-the-routing-table-is-built/how-cef-works)
    * [NSF – Non Stop Forwarding](/ipv4/ipv4-routing/how-the-routing-table-is-built/nsf-non-stop-forwarding)
    * [Routing Order of Operations](/ipv4/ipv4-routing/how-the-routing-table-is-built/routing-order-of-operations)
  * [PfR 101 – Perfromance Routing](/ipv4/ipv4-routing/pfr-101-perfromance-routing)
  * [RIP](/ipv4/ipv4-routing/rip)
    * [RIP 101](/ipv4/ipv4-routing/rip/rip-101)
  * [EIGRP](/ipv4/ipv4-routing/eigrp)
    * [EIGRP 101](/ipv4/ipv4-routing/eigrp/eigrp-101)
    * [EIGRP Metric](/ipv4/ipv4-routing/eigrp/eigrp-metric)
    * [More EIGRP Features](/ipv4/ipv4-routing/eigrp/more-eigrp-features)
  * [OSPF](/ipv4/ipv4-routing/ospf)
    * [OSPF 101](/ipv4/ipv4-routing/ospf/ospf-101)
    * [OSPF Areas](/ipv4/ipv4-routing/ospf/ospf-areas)
    * [OSPF LSAs](/ipv4/ipv4-routing/ospf/ospf-lsas)
    * [OSPF Mechanics](/ipv4/ipv4-routing/ospf/ospf-mechanics)
  * [BGP](/ipv4/ipv4-routing/bgp)
    * [BGP 101](/ipv4/ipv4-routing/bgp/bgp-101)
    * [BGP Attributes](/ipv4/ipv4-routing/bgp/bgp-attributes)
    * [More BGP](/ipv4/ipv4-routing/bgp/more-bgp)
  * [Route Redistribution](/ipv4/ipv4-routing/route-redistribution)
  * [IS-IS](/ipv4/ipv4-routing/is-is)
    * [IS-IS 101](/ipv4/ipv4-routing/is-is/is-is-101)
    * [IS-IS Mechanics – CLNP](/ipv4/ipv4-routing/is-is/is-is-mechanics-clnp)

### IPv6

* [IPv6-101](/ipv6/ipv6-101)
* [IPv6 Routing](/ipv6/ipv6-routing)
* [Interconnecting IPv6 and IPv4](/ipv6/interconnecting-ipv6-and-ipv4)

### MPLS

* [MPLS 101](/mpls/mpls-101)
* [MPLS L3 VPN](/mpls/mpls-l3-vpn)

### Multicast

* [Multicast 101](/multicast/multicast-101)
* [PIM 101](/multicast/pim-101)
* [IGMP 101](/multicast/igmp-101)
* [Inter Domain Multicast](/multicast/inter-domain-multicast)
* [IPv6 Multicast](/multicast/ipv6-multicast)
* [Multicast features on switches](/multicast/multicast-features-on-switches)

### Security

* [NAT 101](/security/nat-101)
* [NAT for Overlapping Networks](/security/nat-for-overlapping-networks)
* [ACLs 101](/security/acls-101)
* [ACLs 102](/security/acls-102)
* [Cisco IOS Firewall](/security/cisco-ios-firewall)
* [Zone Based Firewall](/security/zone-based-firewall)
* [AAA 101](/security/aaa-101)
* [Controlling CLI Access](/security/controlling-cli-access)
* [Control Plane](/security/control-plane)
* [Switch Security](/security/switch-security)
  * [Switchport Traffic Control](/security/switch-security/switchport-traffic-control)
  * [Switchport Port Security](/security/switch-security/switchport-port-security)
  * [DHCP Snooping and DAI](/security/switch-security/dhcp-snooping-and-dai)
  * [802.1x](/security/switch-security/802.1x)
  * [Switch ACLs](/security/switch-security/switch-acls)
* [IPSec VPN 101](/security/ipsec-vpn-101)
  * [IPSEC Crypto Maps 101](/security/ipsec-vpn-101/ipsec-crypto-maps-101)
  * [IPSEC VTI 101](/security/ipsec-vpn-101/ipsec-vti-101)
  * [DMVPN 101](/security/ipsec-vpn-101/dmvpn-101)

### Network Services

* [NTP 101](/network-services/ntp-101)
* [HTTP 101](/network-services/http-101)
* [File Transfer 101 – TFTP & FTP](/network-services/file-transfer-101-tftp-and-ftp)
* [WCCP 101](/network-services/wccp-101)

### QoS

* [QoS 101](/qos/qos-101)
* [Classification and Marking](/qos/classification-and-marking)
* [Congestion Management](/qos/congestion-management)
  * [Legacy Congestion Management](/qos/congestion-management/legacy-congestion-management)
  * [SPD – Selective Packet Discard](/qos/congestion-management/spd-selective-packet-discard)
  * [CBWFQ](/qos/congestion-management/cbwfq)
  * [IP RTP Priority](/qos/congestion-management/ip-rtp-priority)
* [Congestion Avoidance – WRED](/qos/congestion-avoidance-wred)
* [Policing and Shaping](/qos/policing-and-shaping)
  * [CAR 101](/qos/policing-and-shaping/car-101)
* [Compression and LFI](/qos/compression-and-lfi)
  * [Header and Payload Compression](/qos/compression-and-lfi/header-and-payload-compression)
  * [LFI for MultiLink PPP](/qos/compression-and-lfi/lfi-for-multilink-ppp)
* [Frame Relay QoS](/qos/frame-relay-qos)
  * [Per VC Frame Relay QoS](/qos/frame-relay-qos/per-vc-frame-relay-qos)
* [RSVP 101](/qos/rsvp-101)
* [Switching QoS](/qos/switching-qos)

### Network Optimization

* [NetFlow 101 – TNF – Traditional NetFlow](/network-optimization/netflow-101-tnf-traditional-netflow)
* [NetFlow 102 – FNF – Flexible NetFlow](/network-optimization/netflow-102-fnf-flexible-netflow)
* [IP SLA 101](/network-optimization/ip-sla-101)
* [IP Accounting 101](/network-optimization/ip-accounting-101)
* [Logging 101](/network-optimization/logging-101)
* [SNMP and RMON 101](/network-optimization/snmp-and-rmon-101)
* [Cisco CLI Tips and Tricks](/network-optimization/cisco-cli-tips-and-tricks)
* [AutoInstall](/network-optimization/autoinstall)
* [Enhanced Object Tracking](/network-optimization/enhanced-object-tracking)


# Ethernet Switching


# L2 Switch Operations

## CAM - Content Addressable Memory

A switch operates by forwarding frames based on the L2 MAC Destination Address. For each frame it receives, the switch looks up the destination address in the mac address-table and will find out on which port the destination is expected to be found.

The mac address-table is usually automatically built by the switch but it can also accept configurations to manipulate the mac address-table.

In order to build the mac address-table the switch uses the L2 MAC Source Address of a frame to update the table, recording the MAC address, the switchport the VLAN for the incoming frame and a timestamp of the arrival time. If the information already exists, only the timestamp is updated.

When looking up a Destination MAC address in the table there are 2 options

* an entry with the MAC address, port and VLAN exists in the table: In this case the frame is forwarded on the port.
* no entry is found: In this case the frame is forwarded to all ports in the same VLAN as the incoming port. This operation is also known as "Unknown Unicast Flooding"

The mac address-table is also known as CAM (Content Addressable Memory). A CAM works differently than a RAM (Random Access Memory). With RAM you can ask for a the content at a specific address, while with CAM you can ask for the address of a specific content.

To verify the contents of the mac address-table, you can use:

```bash
Sw# show mac address-table [interface INTF]
```

The mac address-table has an aging time. Each entry is kept in the table for until the aging time expires. By default this is set to 300 seconds (5 minutes) but it can be changed in config mode:

```bash
Sw# show mac address-table aging-time
Global Aging Time:  300
Vlan    Aging Time
----    ----------
Sw# conf t
Sw(config)# mac address-table aging-time SEC
# Use 0 to disable aging
```

## TCAM - Ternary Content Addressable Memory

Ternary CAM means this memory supports a third state as well, besides 0 and 1. The third state is X="don't care". This is implemented through a VMR format (Value, Mask, Result)&#x20;

The TCAM is used to hold security ACLs and QoS ACLs and frames would be tested against the TCAM entries to see if the frame should be sent or with what piority. TCAM is also used for L3 forwarding.&#x20;

The allocation of memory for TCAM tables is limited and statically allocated during the boot process but it can be slightly tweaked to use one of the possible SDM (Switching Database Manager) Templates.

```
Sw# show sdm prefer
...
Sw(config)# sdm prefer {vlan|advanced}
```

A reload will be required for the SDM templates to take effect.


# Spanning Tree

{% content-ref url="/pages/19dWvIXBrCDwonlLzAlJ" %}
[802.1d – STP](/layer-2-technologies/ethernet-switching/spanning-tree/802.1d-stp)
{% endcontent-ref %}

{% content-ref url="/pages/jjIc7pGeBl8m6zeaSmeW" %}
[802.1w – RSTP](/layer-2-technologies/ethernet-switching/spanning-tree/802.1w-rstp)
{% endcontent-ref %}

{% content-ref url="/pages/BAKVWjbpgT17SgCOKy48" %}
[802.1s – MSTP](/layer-2-technologies/ethernet-switching/spanning-tree/802.1s-mstp)
{% endcontent-ref %}


# 802.1d – STP

## BPDUs – Bridge Protocol Data Units

BPDUs are sent as 802.3 frames with the source address set to the MAC address of the sending switch port and with a destination address of **01:80:C2:00:00:00**. The LLC header has the DSAP and SSAP fields set to 0x42.

![](/files/011wmklplwu6qeYQGyUk)

* **Protocol ID** is set to 0x0000 (IEEE 802.1d)
* The **version** can be one of 0x00 (802.1d STP), 0x02 (802.1w – RSTP) and 0x03 (802.1s – MSTP).
* The **Type** field indicates wether this BPDU is a Config BPDU(0x00) or a TCN – Topology Change Notification BPDU (0x80).
* Bit 1 in the **Flags** field signals a TCN BPDU, while Bit 8 in the same field signals a TC ACK.
* Timer fields like **Message Age**, **Max Age**, **Hello Time** and **Forward Delay** are stored in units of 1/256 sec.

## How a Bridge ID is unique in a VLAN

To find out the Base Mac Address used by STP, you can use the following command:

```
Sw# show version | i Mac
```

Also, to see Spanning Tree parameters used for each vlan, you can use:

```
Sw# show spanning-tree bridge
                                                   Hello  Max  Fwd
Vlan                         Bridge ID              Time  Age  Dly  Protocol
---------------- --------------------------------- -----  ---  ---  --------
VLAN0001         32769 (32768,   1) 000d.edec.9680    2    20   15  rstp
VLAN0010         32778 (32768,  10) 000d.edec.9680    2    20   15  rstp
VLAN0028         32796 (32768,  28) 000d.edec.9680    2    20   15  rstp
```

### Traditional

This method is used by default if the switch can support 1024 unique MAC addresses for its own use. STP uses one MAC address for each instance and since the MAC address is unique, the Bridge ID becomes unique.

![](/files/CnWVpDDg5daD8oS5Hrtz)

Bridge Priority: 0 – 65535. Default: 32768

### Extended

This method is used if the switch can’t use 1024 unique MAC addresses for itself. STP uses one MAC address for all instances, but uses the VLAN-ID to make the Bridge ID unique for each VLAN.

![](/files/k1QrHjxaozfgNltSXtXP)

The Bridge ID is split into Priority Multiplier (4 bits) and VLAN-ID (12 bits). This means that the Bridge Priority is equal to 4096k + VLAN\_ID. (0<=k<=15). Default BridgeID = 32768 + VLAN\_ID, K=8. The use of the Extended System ID can be disabled, using:

```
Sw(config)# no spanning-tree extended system-id
```

### Setting the Bridge ID

The Bridge ID can be set manually or by using a macro:

```
! Manual:
Sw(config)# spanning-tree vlan VLAN-LIST priority BRIDGE-PRIORITY
! Macro:
Sw(config)# spanning-tree vlan VLAN-LIST root {primary | secondary} [diameter DIA]
```

The macro command will use the following algorithm to determine what BRIDGE-PRIORITY to assign the switch:

* For **primary**:
  * If the current Root Bridge Priority is more than 24576 (k=6), then set the BRIDGE-PRIORITY to 24756
  * If the current Root Bridge Priority is less than or equal to 24576 then set the BRIDGE-PRIORITY to a value 4096 less than the current Root Bridge Priority
  * If the BRIDGE-PRIORITY should be set to less than 4096, then the macro command will fail and the priority should be manually set
* For **secondary**:
  * Sets the BRIDGE-PRIORITY to 28762 (k=7)

Using the root macro will overwrite any custom STP timers that were configured

## Port Cost

In early implementations, a linear scale is used for the cost of a link:

$$
Cost = \frac{1000 Mbps}{Port Bandwidth}
$$

The disadvantage of using such a formula is that links with bandwidth higher than 1 Gbps (1000 Mbps) were considered of equal cost (1).\
Nowadays, a non-linear scale is used, according to the following table:

| Bandwidth | STP Cost |
| --------- | -------- |
| 4 Mbps    | 250      |
| 10 Mbps   | 100      |
| 16 Mbps   | 62       |
| 45 Mbps   | 39       |
| 100 Mbps  | 19       |
| 155 Mbps  | 14       |
| 622 Mbps  | 6        |
| 1 Gbps    | 4        |
| 10 Gbps   | 2        |

To modify the default port cost, use:

```
Sw(config-if)# spanning-tree [vlan VLAN-LIST] cost COST
! cost: 1-65535
! the cost can be set for all vlans, or just for the ones in the VLAN-LIST
```

To verify, use:

```
Sw# show spanning-tree interface INTERFACE [cost]
Vlan             Role Sts Cost      Prio.Nbr Type
---------------- ---- --- --------- -------- --------------------------------
VLAN0010         Desg FWD 19        128.48   P2p
VLAN0028         Desg FWD 19        128.48   P2p
VLAN0030         Desg FWD 19        128.48   P2p
```

Setting the port cost is a method of influencing which ports are in FWD and which are in BLK state on a switch. Setting the cost on the downstream switch will influence the port election on that switch.

## Port ID

The Port ID is a combination of Port Priority and Port Number. The Port Priority can be set as multiples of 16 in the range 0-255. Default value is 128. The Port Number is also in the range 0-255. You cannot change the port Number, but you can change its priority, using:

```
Sw(config-if)# spanning-tree [vlan VLAN-LIST] port-priority PORT-PRIORITY
```

To verify, use:

```
Sw# show spanning-tree interface INTERFACE
Vlan             Role Sts Cost      Prio.Nbr Type
---------------- ---- --- --------- -------- --------------------------------
VLAN0010         Desg FWD 19        128.48   P2p
VLAN0028         Desg FWD 19        128.48   P2p
VLAN0030         Desg FWD 19        128.48   P2p
```

To see the Upstream port priority, look for ‘designated port id X.X’ when using:

```
Sw# show spanning-tree [vlan VLAN-ID] [interface INTERFACE] detail
Port 48 (FastEthernet0/48) of VLAN0010 is forwarding
Port path cost 19, Port priority 128, Port Identifier 128.48.
Designated root has priority 32778, address 000a.f472.f380
Designated bridge has priority 32778, address 000d.edda.eb80
Designated port id is 128.48, designated path cost 4
Timers: message age 0, forward delay 0, hold 0
Number of transitions to forwarding state: 1
Link type is point-to-point by default
BPDU: sent 69469, received 5
```

A lower port priority is preferred. This method is used to influence the election of a forwarding port on a downstream switch, by setting a lower priority for a port on the upstream switch.

## Determining the Best BPDU

The best BPDU is chosen using the following algorithm:

1. Lowest Root Bridge ID
2. Lowest Path Cost to the Root Bridge
3. Lowest Sender Bridge ID
4. Lowest Sender Port ID
5. Lowest Receiver Port ID – not part of the BPDU, but evaluated locally

The BPDU mechanism works as follows:

1. When starting up, all switches think they are the Root Bridge. In 802.1d, as long as a switch thinks it is the Root Bridge, it will generate Config BPDUs every HelloTime and will send them out all (designated) ports, towards the other switches
2. If a switch receives a BPDU with a lower ID in the Root Bridge field, it will not consider itself anymore as the Root Bridge and will stop sending config BPDUs with itself as Root. It will only send out the Designated ports a modified version of the Config BPDUs received from the Root Bridge (on the Root Port).
3. A switch will not send BPDUs out on a port as long as it receives better BPDUs than the ones it would out that port. If better BPDUs are not received for a period of time (MaxAge = 20 sec default), the local port can once again send BPDUs.
4. As an exception, if a switch receives an inferior BPDU on a Designated Port, it will send a BPDU out that port to update the other switch

## STP Convergence

The STP Convergence is achieved in 3 steps:

1. Elect a Root Bridge
2. Elect Root Ports
3. Elect Designated Ports

### Electing a Root Bridge

When they first boot, all switches start sending BPDUs with themselves as the Root, until they receive a BPDU with a lower Bridge ID. In the end, all switches will receive BPDUs with the lowest Bridge ID as the Root Bridge. This Bridge will be considered the Root Bridge.

### Port Roles

To see the role of a port in Spanning Tree, use:

```
Sw# show spanning-tree interface INTERFACE
```

#### **Root Ports**

The Root port is the port on the switch that has the best path to the Root Bridge. Unless the switch considers itself to be the Root Bridge, each switch will elect one port as the Root Port for each STP instance. The port where the best BPDU is received will be the Root Port. The Root Ports will transit to the forwarding state – FWD

#### **Designated Ports**

On each segment, there are usually 2 or more switches. At first, each switch will think it is the designated switch on the segment, and will forward BPDUs on the segment. Eventually, all switches on the segment will find which is the switch that sends the best BPDUs on the segment. This switch will become the Designated Bridge and the port that sends the best BPDU on the segment will become the Designated Port.\
All ports on the Root Bridge should become Designated Ports, unless there is a loop connecting 2 ports on the Root Bridge.\
The Designated ports will transit to the forwarding state – FWD

#### **Backup Ports**

All non-Root and non-Designated ports become Backup Ports.\
All Backup ports will transit to the Blocked State – BLK.

### Port States

In 802.1d, there are 4 port states in addition to the Disabled state where the port is shut down or not participating in the STP process.

* Blocking State – BLK
* Listening State – LIS
* Learning State – LRN
* Forwarding State – FWD

To transit from BLK to FWD, one port must go through the LIS and LRN phase.\
Ports can directly transition from all other states to the BLK state.\
To see the state of a port in spanning tree, use:

```
Sw# show spanning-tree interface INTERFACE
Sw# debug spanning-tree switch state
```

#### **Blocking State – BLK**

All ports start in the BLK State. Ports can transition to the LIS state if they are elected to become Designated or Root ports. This is done only in the following situations:

* if the port does not receive better BPDUs than the ones it would send for a MaxAge period (Default: 20 sec)
* immediately in case of direct failures of the local root/designated port, based on port priority rules

Ports in the BLK state can only receive BPDUs. Receiving no BPDUs in a MaxAge period will transition the port to the LIS state. All Backup ports are immediately put in the BLK state.

#### **Listening State – LIS**

Once in the LIS State, a port can both send and receive BPDUs, but it cannot forward data. Ports in the LIS state will transition back to the BLK state if they lose their Designated or Root status. This can only happen if they receive better BPDUs.\
Ports in the LIS state will transition to the LRN state if they don’t receive better BPDUs in a Forward Delay period (Default: 15 sec).

#### **Learning State – LRN**

Once in the LRN state, a port can both send and receive BPDUs, it still cannot forward data, but it will build its MAC address table based on the traffic it receives. Ports in the LRN state will transition back to the Blocking state if they lose their Designated or Root status. This can only happen if they receive better BPDUs.\
Ports in the LRN state will transition to the FWD state if they don’t receive better BPDUs in a Forward Delay period (Default: 15 sec).

#### **Forwarding State – FWD**

Root and Designated ports end up in the forwarding state. This state permits the port to send and receive BPDUs, but also to forward traffic (finally). A port in the Forwarding State can only move to the blocking state if it becomes a Backup port (loses its Designated or Root status).\
When a port first initializes it will take up to 50 seconds for it to get to the FWD state:

$$
MaxAge\_BLK+FwdDelay\_LIS+FwdDelay\_LRN = 20+15+15 = 50 sec
$$

```
! Example on a Root Bridge
SW1 #show spanning-tree [brief] [vlan VLAN-ID]
VLAN1
  Spanning tree enabled protocol ieee
  Root ID    Priority    32768
             Address     aaaa.aaaa.aaaa
             This bridge is the root
             Hello Time   2 sec  Max Age 20 sec  Forward Delay 15 sec

  Bridge ID  Priority    32768
             Address     aaaa.aaaa.aaaa
             Hello Time   2 sec  Max Age 20 sec  Forward Delay 15 sec
             Aging Time 300

Interface                                   Designated
Name                 Port ID Prio Cost  Sts Cost  Bridge ID            Port ID
-------------------- ------- ---- ----- --- ----- -------------------- -------
FastEthernet1/13     128.54   128    19 FWD     0 32768 aaaa.aaaa.aaaa 128.54
FastEthernet1/15     128.56   128    19 FWD     0 32768 aaaa.aaaa.aaaa 128.56

! Example on a non-root Bridge:
SW2#show spanning-tree [brief] [vlan VLAN-ID]

VLAN1
  Spanning tree enabled protocol ieee
  Root ID    Priority    32768
             Address     aaaa.aaaa.aaaa
             Cost        19
             Port        54 (FastEthernet1/13)
             Hello Time   2 sec  Max Age 20 sec  Forward Delay 15 sec

  Bridge ID  Priority    32768
             Address     cccc.cccc.cccc
             Hello Time   2 sec  Max Age 20 sec  Forward Delay 15 sec
             Aging Time 300

Interface                                   Designated
Name                 Port ID Prio Cost  Sts Cost  Bridge ID            Port ID
-------------------- ------- ---- ----- --- ----- -------------------- -------
FastEthernet1/13     128.54   128    19 FWD     0 32768 aaaa.aaaa.aaaa 128.54
FastEthernet1/15     128.56   128    19 BLK    19 32768 bbbb.bbbb.bbbb 128.55
```

## TCN – Topology Change Notification

There are 2 types of BPDUs.

* Config BPDUs are generated by the Root Bridge and go downstream to the Spanning Tree’s leafs.
* TCN BPDUs go upstream towards the Root Bridge and are the only BPDUs sent out the Root Port.

<figure><img src="/files/2DlAxDzIWs0Dvcaug4VU" alt=""><figcaption><p>TCN Messages sequence</p></figcaption></figure>

How TCN BPDUs work:

1. A bridge originates a TCN BPDU if:
   * It transitions a port into FWD state and it has at least one Designated Port
   * It transitions a port form FWD or LRN to BLK state
2. A bridge will send TCN BPDUs every Hello Time (locally configured) until the TCN is acknowledged.
3. The upstream bridge sets the TCN ACK flag in it’s next BPDU that it sends downstream to acknowledge the TCN BPDU received.
4. The upstream bridge propagates the TCN BPDU out its Root Port.
5. When the TCN BPDU arrives at the Root Bridge, the Root Bridge sets the TCN Acknowledgement flag and the Topology Change flag in the next Configuration BPDU.
6. These flags instructs all bridges to shorten their bridge table aging process from the default 300 sec to the current FwdDelay, thus reducing the convergence time from 300 sec to 50 sec.
7. The Root Bridge continues to set the TCN flag in all its Config BPDUs for the total of FwdDealy + MaxAge (default:35 sec)
8. Downstream bridges will forward the TC BPDUs to other bridges

### STP Timers

Timers can only be modified on the Root Bridge. Timer values are carried in Config BPDUs. Modifications on other bridges will be ignored until they become Root Bridge.

### Hello Time

HelloTime is the time between Config BPDUs sent by the Root Bridge. It is also the interval between TCN BPDUs until the bridge receives a TCN ACK. Default: 2 sec. Lowering Hello Time to 1 sec does not improve convergence\
Hello Time can be set manually:

```
Sw(config)# spanning-tree [vlan VLAN-LIST] hello-time SEC
!hello_time: 1 - 10, default: 2
```

or by using a macro:

```
Sw(config)# spanning-tree vlan VLAN-LIST root {primary | secondary} [diameter DIAMETER] [hello-time SEC]
```

### Forward Delay

FwdDelay is the duration of Listening and Learning States. Also, when a bridge receives a BPDUs with TC flag set, it will set the bridge table aging period to FwdDelay. Default: 15 sec (This value was calculated using a max radius of 7 bridges, a Hello Time of 2 sec and a max of 3 lost BPDUs).\
RFC dictates that time spent in the LIS and LRN states must be equal.\
A simplified formula for FwdDelay calculation is:

$$
FwdDelay = \frac{4*HELLO + 3*DIAMTER}{2}
$$

For details, see [Cisco’s article on Understanding and Tuning Spanning Tree Protocol Timers](https://www.cisco.com/en/US/tech/tk389/tk621/technologies_tech_note09186a0080094954.shtml)

Default FwdDelay is 15 sec. You can change it with:

```
Sw(config)# spanning-tree [vlan VLAN-LIST] forward-delay SEC
!SEC: 4 - 30, default: 15
```

### Max Age

How much time is the Best BPDU stored. If the Best BPDU is not received anymore, it is aged out, and a new Root/Designated Port Election will take place. Convergence is achieved in maximum 50 sec(20+15+15). Direct failures can make the port skip the Max Age timer and put another port directly into a Listening State. In the latter case, convergence is achieved in 30 sec(15+15)\
A simplified formula for Max Age is:

$$
]MaxAge = 4*HELLO + 2*DIAMETER – 2
$$

For details, see [Cisco’s article on Understanding and Tuning Spanning Tree Protocol Timers](https://www.cisco.com/en/US/tech/tk389/tk621/technologies_tech_note09186a0080094954.shtml)

The default MaxAge value is 20 sec. To change it, use:

```
Sw(config)# spanning-tree [vlan VLAN-LIST] max-age SEC
! SEC: 6 - 40, default: 20
```

### Message Age

Message Age is set to 0 on BPDUs originated by the Root Bridge. Bridges that process BPDUs will increment this value by 1 at every bridge hop before forwarding the BPDU.\
If BPDUs from the Root are missed, the switch will count the missed BPDUs and when a BPDU from the root is finally received, it will increment this value once for each missed BPDU.

### Bridge Table Aging Time

When a BPDU with TCN is received, the Bridge Table Aging Time will change from 300 sec to FwdDelay to reduce convergence time.

## Cisco STP Enhancements

### Portfast

This feature enables the transition of a port from the BLK state directly to the FWD state, bypassing LIS and LRN states. This feature should only be used on access ports where only one host is connected.\
Transitions from FWD to BLK or BLK to FWD on ports configured as portfast do not generate TCN BPDUs out the Root Port. Ports configured as Portfast still run spanning tree in order to block a potential loop, so they still send BPDUs.\
[BPDU Guard](https://nyquist.eu/802-1d-stp/#102_BPDU_Guard) and [BPDU Filter](https://nyquist.eu/802-1d-stp/#103_BPDU_Filter) can be enabled on portfast ports.\
To enable the portfast status of a port, you can use the command:

```
Sw(config-if)# spanning-tree porfast
```

You can also use a macro that enables portfast, sets the port to access mode and disables PAgP

```
Sw(config-if)# switchport host
```

The last option is to use a global command that will enable portfast for all non-trunking ports:

```
Sw(config)# spanning-tree portfast default
```

To verify, use:

```
Sw# show spanning-tree interface INTERFACE portfast
```

When portfast is off, convergence is achieved in 50 sec, and when it is on, it is instant.

### Uplink Fast

If a switch has multiple uplinks that lead to the Root Bridge, only one of them will be elected as the Root Port, and all others will be Backup Ports. When the primary uplink fails, a new one will be elected as the Root Port, but it still has to go from BLK through LIS and LRN states, before reaching the FWD state.\
The Uplink Fast feature enables leaf-mode switches at the end of the spanning tree branches to have a functioning Root Port while keeping one or more redundant or potential Root Ports in Blocking state. When the primary root port uplink fails, the next lowest Root Path Cost port is unblocked and used without delay.\
Uplink Fast is enabled per switch, on all VLANs. To enable it, use:

```
Sw(config)# spanning-tree uplink-fast [max-update-rate PPS]
!PPS: 0-65535. Default: 150
```

The command will not be allowed on the Root Bridge. When entering this command the switch will automatically:

* **Raise Bridge Priority to 49152** (k=12) to make sure that this switch does not become the Root Bridge and that the switch is not used as a transit switch to get to the Root Bridge
* **Increment Port Cost by 3000** making the ports undesirable as paths to the root for any downstream switches.

If the root port fails and uplink fast is enabled, as soon as a new root port is elected, the switch will sends multicast frames to 0100.0CCD.CDCD on behalf of all entries in the CAM table so that upstream switches can learn about the new uplink. These frames are sent at a configurable rate (max-update-rate). If the rate is set to 0, these frames are not sent.\
To verify:

```
Sw# show spanning-tree uplink-fast
```

Uplink Fast reduces convergence in case of direct failures from 30-50 sec down to zero.

### Backbone Fast

It helps STP to converge faster, in case of indirect link failure.\
A switch detects an indirect link failure when it receives inferior BPDUs from the Designated Bridge on a Root Port or a Blocked Port. When this happens, the switch will use the RLQ(Root Link Query) protocol to see if upstream switches have a stable path to the Root Bridge. RLQ is enabled only when Backbone Fast is enabled, so you must use Backbone Fast on all switches in the network:

```
Sw(config)# spanning-tree backbone-fast
```

To verify:

```
Sw# show spanning-tree backbone-fast
```

If backbone fast is enabled, when the switch receives an inferior BPDU:

* If the inferior BPDU came on a Blocked Port – The switch considers the Root Port and all other blocked ports as alternate paths to the Root Bridge. It starts using RLQ
* If the inferior BPDU came on the Root Port
  * If there are no Blocked Ports – The switch assumes it is the Root Bridge – This is done quicker than by using the default MaxAge mechanism
  * If there are Blocked Ports – The switch considers all Blocked Ports as alternate paths to the Root Brige. It starts using RLQ

#### **RLQ – Root Link Queries**

RLQ Requests are sent out all alternate paths.\
When a switch receives a RLQ request it will send a RLQ reply if it thinks it is the Root Bridge (it actually is or it lost connectivity to the Root Bridge). In any other case, it will relay the RLQ request to all other connected switches.\
The switch that send the RLQ Request will eventually receive a RLQ Reply from the Root Bridge.

* If the reply came on the current Root Port, than the path to the Root is stable.
* If the reply came on a non-Root Port, then an alternate path must be used. MaxAge will immediately expire so the port will go directly to the LIS state and from there to LRN and FWD. This way, convergence in case of indirect failure is reduced from 50 to 30 sec.

## Protection Mechanisms

### Root Guard

Root Guard should be used on ports where you do not expect root bridges to appear. It will protect against better BPDUs on ports where the Root Bridge should not be found.\
If better BPDUs are received on a Root Guard enabled port, the switch will move the port in a “Root-inconsistent” state. In this state, no data is sent or received, but the switch can listen to BPDUs. If BPDUs are not received anymore, the port enters normal STP cycle.

```
Sw(config-if)# spanning-tree guard root
Sw# show spanning-tree inconsistent-ports
Name                 Interface              Inconsistency
-------------------- ---------------------- ------------------
VLAN0021             FastEthernet0/1            Root Inconsistent

Number of inconsistent ports (segments) in the system : 1
```

Another indication can be seen while looking at the spanning tree data for the VLAN:

```
Sw# show spanning-tree vlan 21 | b Interface
Interface           Role Sts Cost      Prio.Nbr Type
------------------- ---- --- --------- -------- --------------------------------
Fa0/1               Desg BKN*100       128.66   Shr *ROOT_Inc 
! The port is blocked with the status ROOT_Inc
```

### BPDU Guard

BPDU Guard is a feature that protects against any BPDUs received. If any BPDUs are received on such ports, they are immediately put into errdisable state.\
In errdisable state, the ports are shutdown and have to be manually re-enabled (shut/no shut) or they have to wait for errdisable to timeout.\
By default, BPDU Guard will shutdown all VLANs on the port, but it can be configured to disable only the offending VLAN if you configure:

```
Sw(config)# errdisable detect cause bpduguard shutdown vlan
```

BPDU Guard can be enabled per port, regardless of the portfast status:

```
Sw(config-if)# spanning-tree bpduguard enable
```

or it can be enabled globally on every portfast port, using:

```
Sw(config)# spanning-tree portfast bpduguard default
```

To enable auto-recovery for bpduguard errordisabled ports, use:

```
Sw(config)# errdisable recovery cause bpduguard
! To configure the auto-recovery interval:
Sw(config)# errdisable recovery interval INTERVAL
! default 300 sec
```

### BPDU Filter

BPDU Filter disables sending and receiving of BPDUs on the ports it is configured on. It is as if STP is not runnning on those ports. BPDU Filtering can be be enabled globally or per port, but there is a difference in how it works:\
To enable BPDU Filtering for all portfast ports:

```
Sw(config)# spanning-tree portfast bpdufilter default
```

This way, all portfast ports will stop sending BPDUs. If a BPDU is received on a PortFast port with BPDU filtering, the port will lose its portfast status and will disable BPDU filtering. It will now work like a normal STP port.

You can also enable BPDU filtering on each port, regardless of the portfast status.

```
Sw(config-if)# spanning-tree bpdufilter {enable | disable}
```

This way, the port acts as if STP is disabled and may lead to loops in the network.

### Loop Guard

When enabled, it keeps track of the BPDU activity on non-designated ports. As long as BPDUs are received, the port is allowed to behave normally. When BPDUs go missing, Loop Guard moves the port into Loop-inconsistent state, so that it cannot go into FWD after MaxAge expires.\
To enable Loop Guard, use:

```
!Per port:
Sw(config-if)# spanning-tree guard loop
!Globally
Sw(config)# spanning-tree loopguard default
```

### UDLD

UDLD is a Cisco proprietary feature that interactively monitors a port to see if the link is truly bidirectional. UDLD expects the far-end switch to echo these frames back across the same link, with the far-end switch port’s identification added. If echoes are not seen back, the link is considered unidirectional.\
There are 2 modes of operation:

* **Normal**: when an uni-directional link is detected, the port is allowed to continue operations. The port is marked in an undetermined state and a syslog message is generated.
* **Aggressive**: when an uni-directional link is detected, the port sends 8 UDLD messages, once a second. If none is echoed back, the port is put in errdisable state

UDLD doesn’t do anything until the other switch echoes back messages. So, unless UDLD worked before, the link won’t be disabled.\
UDLD can be enabled globally for all optic links, using:

```
Sw(config)# udld {enable|aggressive}
```

or it can be enabled on each port, regardless of type, using:

```
Sw(config-if)# udld {enable|aggressive|disable}
```

The time between probe messages can be adjusted globally with:

```
Sw(config)# udld message time SEC
```

To verify, use:

```
Sw# show udld INTERFACE
```

To reset all ports blocked by UDLD, use:

```
Sw# udld reset
```

## IEEE, PVST and PVST+

Originally, IEEE adopted the 802.1d Spanning Tree standard. BPDUs were carried over 802.1q links as untagged frames. This means that only one STP instance was supported. This was known as the MST (Mono Spanning Tree) or CST (Common Spanning Tree)

Cisco decided to run one 802.1d instance on each VLAN, resulting in PVST (Per VLAN Spanning Tree). PVST runs only on ISL trunks so it was not inter-operable with IEEE 802.1d

In the end, Cisco came up with PVST+ a version of PVST that runs also on 802.1q links and is interoperable with IEEE 802.1d.

PVST+ sends different BPDUs for each VLAN. The IEEE 802.1d version sends one BPDU regardless of the number of VLANs. When a Cisco switch connects to a IEEE region, it will exchange untagged BPDUs over the native VLAN with the CST instance and will encapsulate BPDUs for the other VLANs inside a multicast frame that will be forwarded by the IEEE switches out all ports, eventually reaching other Cisco Switches that will understand their special meaning. This mode of tunneling is possible because the VLAN ID information is added to the frame as a TLV field.

PVST+ is the default spanning tree mode running on Cisco switches. If MST or rapid-pvst was enabled, you can re-enable PVST+ using:

```
Sw(config)# spanning-tree mode pvst
```


# 802.1w – RSTP

Rapid Spanning Tree is an updated version of the original Spanning Tree protocol, standardized as 802.1w. It includes many of the Cisco proprietary features of PVST+. RSTP is backwards compatible with 802.1d STP but it has a few new features.\
To enable RSTP, use:

```
Sw(config)#spanning-tree mode rapid-pvst
```

## New features in RSTP

### RSTP Port States

RSTP uses only 3 states:

* **Discarding** – Port is not forwarding data so it breaks the loops.
  * Thi state is seen in a stable active topology and during topology synchronization
* **Learning** – Port is in a transition state from DIS to FWD. The learning state accepts data frames to populate the MAC table to limit flooding of unknown unicast frames.&#x20;
  * This state is seen in a stable actice topology an during topology synchronization
* **Forwarding** – Port is forwarding data.&#x20;
  * This state is seen only in a stable active topology

### RSTP Port Roles

* **Root Port** – Same as in 802.1d – The port that receives the best BPDU on a switch is the Root Port. It is considered the closest port to the Root Bridge
* **Designated Port** – Same as in 802.1d – On any segment, the port that sends the best BPDU is considered the Desingated Port.
* **Alternate Port** – Alternate to a Root Port – BPDUs from the Root Bridge can be received on other ports than the Root Port, but they are normally disabled. Still, they can offer an alternate path to the Root Bridge, in case the Root Port fails. This is similar to the UplinkFast feature of PVST+. When the path through the Root Port fails, the switch can quickly chose one of the Alternate Ports as the new Root Port.
* **Backup Port** – Backup of a Designated Port – On a segment, ports on the same switch as the Designated Bridge can be used as Backup ports of the Designated Port. This means that they can offer a backup path for that segment, but they cannot guarantee a different path towards the Root Bridge

### RSTP BPDU Format

802.1D BPDUs are sent with Version set to 1 and Type set to 1.\
RSTP BPDUs are sent with Version set to 2 and Type set to 2. Legacy bridges will drop these frames. Alos, new flags are implemented in the Flags Byte, that encode the Role and State of the sending port:

| Bit       | Function                                                                       |
| --------- | ------------------------------------------------------------------------------ |
| 0         | Topology Change                                                                |
| 1 – NEW   | Proposal                                                                       |
| 2,3 – NEW | <p>00 – Unknown<br>01 – Alternate / Backup<br>10 – Root<br>11 – Designated</p> |
| 4 – NEW   | Learning                                                                       |
| 5 – NEW   | Forwarding                                                                     |
| 6 – NEW   | Agreement                                                                      |
| 7         | Topology Change ACK                                                            |

### New BPDU Handling

In 802.1d, a non-root bridge would only generate BPDUs when it received one on its root port. In 802.1w, a bridge sends BPDUs every Hello Time (default: 2sec), regardless of the BPDUs it receives from the root bridge.\
In 802.1d, BPDUs had to go missing for Max Age(default: 20 sec) before the BPDU information was aged out for that port and new information would be processed. In 802.1w, the BPDUs act as a keepalive mechanism, so if 3 BPDUs are missed, the bridge considers the connection to its neighbor down and it ages out BPDU information for that port.

### Uses a Backbone Fast mechanism by default

RSTP includes a mechanism similar to the [Backbone Fast enhancement in STP](https://nyquist.eu/802-1d-stp/#93_Backbone_Fast) that will work by default without any additional configuration.

### Rapid Transitions to Forwarding State

Rapid Transition is achieved on Edge Ports and on Point to Point links.

#### **Edge Ports**

* An Edge Port is similar to a PVST+ PortFast port.
* An Edge Port does not generate TCN when its link flaps.
* An Edge Port that receives a BPDU loses its Edge status

To set an Edge Port, use the same command as in PVST+:

```
Sw(config-if)# spanning-tree portfast
```

#### **Link Type**

All Links fall in 2 categories in RSTP. By default, all full duplex links are considered Point to Point, while all Half Duplex links are considered Shared Links.\
They can also be manually set, using:

```
Sw(config-if)#spanning-tree link-type {shared|point-to-point}
```

To verify the link type, use:

```
Sw# show spanning-tree [vlan VLAN-ID]
Fa0/1               Desg FWD 19        128.1    P2p
Fa0/2               Desg FWD 19        128.2    Shr
Fa0/3               Desg FWD 19        128.3    P2p Edge
```

### New TCN Mechanism

In RSTP, only non-Edge ports that are moving to a FWD state will cause a TCN. When a bridge sends a TCN, it first starts a TC While timer, equal to 2xHello Time. Over this period, it sends TCN BPDUs out all its non-Edge Designated and Root Ports. It also flushes the MAC addresses associated with these ports.\
When a bridge receives a BPDU with TC bit set, it starts a TC While timer, equal to 2xHello Time. Over this period, it sends TCN BPDUs out all its non-Edge Designated and Root Ports, except the port where it received the TCN. It also flushes the MAC addresses associated with these ports.\
This way, the TCN is quickly flood and there is no need for the TCN to travel upstream to the Root and then downstream, as in 802.1D STP.

## Proposal/Handshake Mechanism

Convergence in RSTP is achieved using a Link-by-link handshake mechanism. On each link, when it comes up, both ports are set to a Designated Role and to a BLK State. The 2 bridges send each other BPDUs with the Proposal bit set. The superior bridge (the Root Bridge or the Designated Bridge on the path to the Root) will transition its designated port to a FWD state only after the inferior bridge will be “in sync”.\
Being “in sync” means the inferior bridge will put all its non-Edge Designated Ports into a BLK state and will start a similar Handshake mechanism Downstream. When all non-Edge Designated Ports are blocked, the first port can transition to a Root Port Role and a FWD State and will send an Agreement BPDU to the superior bridge. The superior bridge can now transition its port into a FWD State. This way the loops are prevented and the “sync” boundary moves downstream to the other bridges.\
As you can see, before agreeing to the RSTP proposal, a bridge must have all other ports “in sync”. A port is considered “in sync” if it is an Edge Port or if it is in a BLK State. If ports are not in BLK, they are transitioned to BLK in order to agree to the proposal.

This kind of mechanism is faster because it doesn’t use timers like in 802.1d STP, but it can only take place over P2P links. On Shared links, the protocol falls back to 802.1d timers mechanism. This also happens if no agreement is received.

## Compatibility with 802.1d STP

802.1d Bridges do not understand RSTP BPDUs, but RSTP bridges understand STP BPDUs. When a RSTP bridge detects an 802.d BPDU, it falls back to using 802.1d mechanism on that port.


# 802.1s – MSTP

The number of STP or RSTP instances that can run on a switch is limited. It is usually enough for most implementations but if you need to run STP on more VLANs, then MSTP is the solution. MSTP maps multiple VLANs to the same STP instance. Even when PVST or Rapid PVST is used, chances are that more than one instance will use the same loop-free topology and this is a waste of resources.

In PVST+ and Rapid PVST, one BPDU is sent on each VLAN on a trunk port. In 802.1q – there was only one STP instance, the CST, so only one BPDU was sent for all VLANs over the 802.1q links.\
To enable MSTP, use:

```
Sw(config)# spanning-tree mode mst
```

## MST Regions

802.1s – MSTP uses the concept of regions. A MST region is represented by all connected switches that have the same VLAN to MSTP instance mapping.\
A MST Region is identified by:

* a 32 Bytes alpahnumeric identifier
* a 2 Bytes revision number
* a table that maps each VLAN to a STP instance

Switches that share the same identification information are part of the same MST Region. When sending BPDUs, switches send a message digest of the VLAN-to-instance table. If the received digest is different then the one computed using its own data, the port where this BPDU is received is considered a Boundary Port.

## MST Instances

Cisco’s implementation of 802.1s supports 16 instances, 1 IST (Internal Spanning Tree) and 15 MSTIs(MST Instances).\
To configure the MST instances, use:

```
Sw(config)# spanning-tree mst configuration
Sw(config-mst)# name MST-DOMAIN-NAME
Sw(config-mst)# revision MST-DOMAIN-REVISION
Sw(config-mst)# instance INSTANCE-ID vlan VLAN-LIST
```

Configuration will only be applied when you exit the mst configuration mode. To verify pending config, use:

```
Sw(config-mst)# show pending
```

When you first enable MSTP, Instance 0 is created and all VLANs are allocated to it.

### IST

The IST runs on all bridges inside a MST Region and also provides interaction with the other MST Regions or other STP flavours (802.1d STP, 802.1q CST, PVST+). IST runs like a RSTP instance, but it also carries timer information in order to interact with legacy STP regions.

### MSTIs

The MSTIs are RSTP instances that exist only within a MST region, with no outside interaction. MSTIs do not send their own BPDUs, instead MSTI information is attached to the IST BPDUs as M-Records. M-Records do not contain timer information so all instances use the same timers as the IST. MSTIs follow the IST topology at boundary port, so a boundary port will FWD or BLK for all VLANs.

## MST Hop Count

MSTP Root Bridges send BPDUs with a default hop count of 20, and each bridge along the way decrements this value. If the hop count reaches 0, the incoming BPDU is discarded. To change the default value, use:

```
Sw(config)# spanning-tree mst max-hops HOP-COUNT
```

## MST Configuration

MST Configuration is similar to the other STP flavours but it uses the mst keyword, like this:

```
! Configuring Root Bridge
Sw(config)# spanning-tree mst INSTANCE-LIST root {primary|secondary} [diameter DIA [hello-time HELLO]]
! Configuring Bridge Priority
Sw(config)# spanning-tree mst INSTANCE-LIST priority PRI
! Configuring Port Priority
Sw(config-if)# spanning-tree mst INSTANCE-LIST port-priority PRI
! Configuring Port Cost
Sw(config-if)# spanning-tree mst INSTANCE-LIST cost COST
```

Timers are only configured for IST, so there is no configuration per-instance:

```
Sw(config)# spanning-tree mst {hello-time HELLO|forward-time FORWARD|max-age MAXAGE}
```

RSTP Link Types are configured just link in RapidPVST

```
Sw(config-if)# spanning-tree link-type {point-to-point|shared}
```

## Interoperability with other STP flavours

MSTP seamlessly interoperates with 802.1q CST.\
When interconnected with a PVST bridge, the MST bridge will replicate the IST out all VLANs. Not all possible configurations will be considered valid and an invalid configuration will put the boundary ports into a Root Inconsistent mode.\
A valid configuration can be achieved in the following cases:

* If the MSTP is the root bridge, it must be the root for all VLANs (recommended)
* If the PVST+ is the root bridge, it must be the root for all VLANs

Any other configuration will fail. (E.g. when the MSTP bridge is the root for the CST while the PVST+ bridge is the root for one or more VLANs)\
You can manually set the neighbor type to non-MSTP, using:

```
Sw(config)#spanning-tree mst pre-standard
```


# VTP 101

## How it works

VTP (VLAN Trunking Protocol) is used to advertise VLAN information between connected switches in a network. VTP messages are sent only on trunk ports. VTP advertises the VLAN\_ID, VLAN\_NAME, VLAN\_TYPE and VLAN\_STATE information for each VLAN. VTP doesn’t send any information regarding port assignment to VLANs.

There are 3 modes VTP can run in: server, client and transparent. Switches running in client mode can’t make changes to the VLAN information and only changes on switches in server mode are advertised to the other switches.

### VTP version 1 and 2

Changes on VTP servers result in the VTP revision number being incremented and the propagation of new VLAN information in VTP messages. Upon receiving the message, a VTP server or client verifies the received revision number against it’s stored revision number and if the received one is higher, will save the new information in the vlan.dat file.

### Improvements in VTP version 3

Due to numerous issues that can occur if there are multiple VTP servers that can change the configuration in a domain, VPTv3 introduces the concept of Primary VTP Server, which is the only switch that can perform changes to the VLAN information (together with the transparent switches, but they are independent). All other VTP servers actually work like clients, the difference being that they can be promoted to a Primary Serer role. To make things safer, switches will only accept VTP changes from the primary server they know and will ignore changes if another primary server comes up, even if its updates have a higher revision number.

## VTP Versions

```
Sw(config)# vtp version {1|2|3}
!Default: 1
```

### Version 1

* supports only Normal Range VLANs: 1-1005
* In transparent mode, it forwards **only version 1** advertisements received on trunk links if VTP domain is NULL or if it matches the domain in the message

### Version 2

* supports only Normal Range VLANs: 1-1005
* In transparent mode, it forwards **any** VTP advertisements received on trunk links if VTP domain is NULL or if it matches the domain in the message

### Version 3

* Supports Normal and Extended Range VLANs (1-4094), including Private VLANs, but is only available starting with IOS 12.2.(52)SE.
* In transparent mode, it forwards VTP advertisements received on trunk links if VTP domain is NULL or if it matches the domain in the message
* Passwords are not shown in clear text anymore.
* Supports an OFF mode where VTP is disabled (per switch, or per interface) and received VTP messages are dropped:

  ```
  Sw(config-if)# no vtp
  ```
* Server role changed to support a single primary VTP server and several secondary VTP servers. Only the primary server can change the VLAN information, and only one can exist at one time. VTP primary server role is requested through an EXEC command
* Backwards compatibile to VTPv1 and v2 by falling back to the detected version on each port
* VTPv3 is a generalized protocol to exchange database information to other switches, so it is able to distribute MST region configurations as well

## VTP Mode

To set the vtp mode, use:

```
Sw(config)# vtp mode {client|server|transparent|off}
! Default: server
! off is only available in VTPv3
```

To see the mode, use:

```
Sw# show vtp status
```

### Server

* It can create, modify and delete VLANs
* It sends and receives VLAN configuration in VTP advertisements, over trunk links.
* It will accept configuration changes from other VTP Servers

In version 3, to make a VTP server primary, use:

```
Sw# vtp primary
```

### Client

* It cannot create, modify and delete VLANs
* It sends and receives VLAN configuration in VTP advertisements, over trunk links.
* It will accept configuration changes from other VTP Servers

### Transparent

* It can create, modify and delete VLANs, but they have local significance
* Forwards VLAN configuration in VTP advertisements, over trunk links, as long as the messages are for the same domain as the switch, or if the switch hasn’t been configured with a domain name (domain=NULL).
* It will not accept configuration changes from other VTP Servers

## Domain

To accept VTP advertisements, switches must be in the same VTP Domain. By default, a switch is set up as a VTP Server, but with a NULL domain name and it will act like a transparent VTP switch. If the VTP Domain is NULL, the switch will set its VTP domain to the first Domain that it sees in a VTP advertisement. You can manually set the domain with:

```
Sw(config)# vtp domain VTP-DOMAIN
```

### Passwords

For each VTP domain, a VTP password can be set in order to authenticate VTP advertisements.

```
Sw(config)# vtp password VTP-PASS [hidden]
! hidden is only available for VTPv3
```

To see the password used:

```
Sw(config)# show vtp password
!It will show either plain text or the encrypted password, depending on how it was configured
```

## VTP Pruning

Pruning prevents VLANs from being carried over trunks where they are not needed. Pruning can be enabled only on one VTP server

```
Sw(config)# vtp pruning
```

VTP Pruning can only prune a list of pruning-eligible VLANs that are configured per interface:

```
Sw(config-if)# switchport trunk pruning vlan {add VLAN-LIST|remove VLAN-LIST|except VLAN-LIST|none}
```


# Private VLANs

Private VLANs partitions a regular VLAN domain into subdomains. Such a subdomain is created when a primary VLAN is paired with a secondary VLAN. Only a switch in VTP Transparent mode supports Private VLANs

## Primary VLAN

To set a VLAN as Primary VLAN, use:

```
Sw(config)#vlan VLAN-ID
Sw(config-vlan)#private-vlan primary
```

After the secondary VLANs are configured, they are associated with the primary VLAN, using the following confing form withing the primary vlan:

```
Sw(config)#private-vlan association [add|remove] SECONDARY-VLAN-LIST
```

To verify, use:

```
Sw# show vlan private-vlan [type]
```

## Secondary VLANs

Secondary VLANs can be condfigured as Isolated or as Community VLANs. Private VLANs work over different switches, as long as the Private VLANs and the primary VLANs are carried over the trunk links.

### Isolated VLANs

Ports within the isolated VLAN cannot communicate with each other at Layer.

```
Sw(config-vlan)# private-vlan isolated
```

### Community VLANs

Ports within the Community VLAN can communicate with ports in the same Community VLAN but not with ports in other Community VLANs or with ports in the Isolated VLAN.

```
Sw(config-vlan)# private-vlan community
```

## Private VLAN Ports

To configure a port as part of a private VLAN, use:

### Promiscuous Ports

A promiscuous ports belongs to the primary VLAN and can communicate with all interfaces in the primary VLAN, including ports in the isolated and community secondary VLANs. To set up a promiscuous port, use:

```
Sw(config-if)# switchport mode private-vlan promiscuous
Sw(config-if)# switchport private-vlan mapping PRI-VLAN-ID {add|remove} SEC-VLAN-LIST
!The promiscuous port must be mapped to one or more Secondary VLANs
```

You can also map the VLANs to a promiscuous port using:

```
Sw(config-if)# switchport private-vlan association mapping PRI-VLAN-ID {add|remove} SEC-VLAN-LIST
```

### Isolated Ports

It is a port that belongs to the Isolted Secondary VLAN. It can only communicate with promiscuous ports.To set an isolated port, use:

```
Sw(config-if)# switchport mode private-vlan host
Sw(config-if)# switchport private-vlan host-association PRI-VLAN-ID SEC-VLAN-ID
!SEC-VLAN-ID must be an isolated VLAN
```

You can alos map the VLANs to a isolated port, using:

```
Sw(config-if)# switchport private-vlan association host PRI-VLAN-ID SEC-VLAN-ID
```

### Community Ports

It is a port that is part of a Community Secondary VLAN and it can only communicate with other ports in the same Community or with promiscuous ports.To set a community port, use:

```
Sw(config-if)# switchport mode private-vlan host
Sw(config-if)# switchport private-vlan host-association PRI-VLAN-ID SEC-VLAN-ID
!SEC-VLAN-ID must be a community VLAN
```

You can alos map the VLANs to a community port, using:

```
Sw(config-if)# switchport private-vlan association host PRI-VLAN-ID SEC-VLAN-ID
```

## Mapping Secondary VLANs to Primary Layer3 VLANs

To allow inter-vlan routing, the secondary VLANs must be mapped to the L3 SVI:

```
Sw(config)# interface vlan PRI-VLAN-ID
Sw(config-if)# private-vlan mapping [add|remove] SEC-VLAN-LIST
```

To monitor, use:

```
Sw# show interface private-vlan mapping
```


# VLANs

## Supported VLANs

Supported VLANs: 1-4094

```
Sw(config)# vlan VLAN-ID
```

Optional parameters:

```
Sw(config-vlan)# name VLAN-NAME
Sw(config-vlan)# mtu MTU-SIZE
!default: 1500
```

To verify:

```
Sw# show vlan [name VLAN-NAME| id VLAN-ID]
```

### Normal Range vs Extended Range

* **Normal Range VLANs**: 1-1005\* – Supported by VTP v1,v2,v3 in all modes.
* **Extended Range VLANs**: 1006-4094 – Only supported in Transparent mode by VTPv1 and v2, and in all modes by VTP v3.

\* Reserved VLANs: 1002-1005 (Token Ring and FDDI)

## Port Modes

### Static Access

```
Sw(config-if)# switchport mode access
```

#### **2.1.1Access with static VLAN assignment**

Default VLAN for all ports is VLAN 1.To change it, use:

```
Sw(config-if)# switchport access vlan VLAN-ID
```

#### **Access with dynamic VLAN assignment**

```
Sw(config-if)# switchport access vlan dynamic
```

This configuration requires a VMPS Server (VLAN Management Policy Server).

### Static Trunk

First, Define the encapsulation type:

```
Sw(config-if)# switchport trunk encapsulation {dot1q|isl|negotiate}
! dot1q =  802.1q IEEE Standard
! isl = Cisco Proprietary (prefered)
! negotiate = used if dynamic trunking is used (default)
```

To configure the port as a static Trunk port, use:

```
Sw(config-if)# switchport mode trunk
```

#### **Native VLANs**

On 802.1q trunks, frames in the native vlan are sent untagged. To set the native vlan, use:

```
Sw(config-if)# switchport trunk native vlan VLAN-ID
! Default Native VLAN: 1.
```

#### **Allowed VLANs**

By default, all VLANs are allowed on a trunk. To limit the VLANs that can travel over a trunk link, use:

```
Sw(config-if)# switchport trunk allowed vlan {add VLAN-LIST| remove VLAN-LIST| except VLAN-LIST|all}
```

### Dynamic

A port configure in dynamic mode, will become either an access or a trunk port, depending on the negotiation with the other pot it is connected to.\
DTP is used to negotiate Dynamic Trunks.\
To configure a port as dynamic, use the following command:

```
Sw(config-if)# switchport mode dynamic {auto|desirable}
!Default setting is platform dependent
```

Here’s how the ports will end up in a dynamic configuration:

| This side\The other side | Access        | Trunk        | Dynamic Auto  | Dynamic Desirable |
| ------------------------ | ------------- | ------------ | ------------- | ----------------- |
| Access                   | access\access | access\trunk | acces\access  | access\access     |
| Trunk                    | trunk\access  | trunk\trunk  | trunk\trunk   | trunk\trunk       |
| Dynamic Auto             | access\access | trunk\trunk  | access\access | trunk\trunk       |
| Dynamic Desirable        | access\access | trunk\trunk  | trunk\trunk   | trunk\trunk       |

DTP works by default, even if the port is configured as a static trunk. This is needed so that the other end of the connection could negotiate to become a trunk.\
DTP is disabled if the port is set in the acccess mode or if the following command is used:

```
Sw(config-if)# switchport nonegotiate
```

This is usually used when a static trunk is created with a neighbor that does not support DTP (like a router or a firewall).\
DTP negociation will fail if the devices are in different VTP domains.

### Private VLAN

See [Private VLANs](/layer-2-technologies/ethernet-switching/private-vlans)

### Dot1Q Tunnels (Q-in-Q Tunnels)

Q-in-Q tunnels are used to carry frames tagged with 802.1q by a customer over a provider’s network. Customer traffic is encapsulated within another 802.1q tag (metro tag) which is used inside the provider network.\
An asymmetric link must be configured, where the port on the customer switch will be set as Trunk port, and the port on the Provider switch will be set as Tunnel Port.

#### **Configure the Customer Switch**

To configure the customer switch, use:

```
SwC(config-if)# switchport mode trunk
```

#### **Configure the Provider Switch**

```
SwP(config-if)# switchport access vlan VLAN-ID
! VLAN-ID is the metro tag, specific to each customer
SwP(config-if)# switchport mode dot1q-tunnel
! Will also disable CDP on the port
```

Since Q-in-Q tunnels add an additional dot1Q header, the MTU of the frames can reach 1504 Bytes. The switch will warn when a port is configured for dot1q tunneling that the default MTU of the switch (1500) should be changed. This change should be done on all provider switches:

```
Sw(config)# system mtu 1504
! Requires a reload
```

To test that the provider network can accomodate such frames, you can try:

```
SwC# ping IP-ADDR size 1500 df-bit
```

#### **Native VLAN on Dot1q Tunnels**

Since trunk belonging to the native VLAN is normally sent untagged, this could end up in problems inside the provider network. To prevent this use one of the following:

* Use ISL encapsulation inside the provider network
* Make sure the native VLAN on the customer/provider edge is not within the customer VLAN range.
* Tag the native VLAN:

  ```
  Sw(config)# vlan dot1q tag native
  ```

#### **Tunneling L2 Protocols**

Normally traffic for VTP, CDP, STP, PAgP, LACP, UDLD is not switched. It is interpreted by each device on a per-link basis.\
To enable L2 Protocl Tunneling, use the following config on the provider switch:

```
SwP(config-if)# l2protocol-tunnel [cdp|stp|vtp]
```

To tunnel protocols used over point-to-point connections, use the following command to emmulate such a connection:

```
SwP(config-if)# l2protocol-tunnel point-to-point {pagp|lacp|udld}
```

The number of L2 packets that can be tunneled can be limited, using:

```
SwP(config-if)# l2protocol-tunnel drop-threshold [cdp|stp|vtp|point-to-point {pagp|lacp|udld}] VALUE
! if no protocol is specified, the limit applies to all protocols
! Packets over the specified threshold will be dropped
```

You can also configure the interface to shutdown if a threshold is violated:

```
SwP(config-if)# l2protocol-tunnel shutdown-threshold [cdp|stp|vtp|point-to-point {pagp|lacp|udld}] VALUE
! if no protocol is specified, the limit applies to all protocols
```

To auto-recover from such an error, use:

```
SwP(config)# errdisable recovery cause l2ptguard
```

To monitor,use:

```
SwP#show l2protocol-tunnel [summary|interface INTERFACE]
```


# EtherChannel 101

An etherchannel is a logical port that consists of multiple links bundled into a single logical link. To have a working etherchannel you must use static config or a negotiation protocol (LACP or PAgP). All ports in an EtherChannel must operate at the same speed and duplex. When an EtherChannel group is first created, it will follow the configuration of the first port that was added to the group. The following configurations must match between all ports in a group:

* Allowed VLAN List and Native VLAN
* Spanning-tree path cost and priority for each VLAN
* Spanning-tree PortFast setting

## Static Configuration

```
Sw(config)# channel-group PO-ID mode on
```

This configuration will work only if all ports in the channel-groups on both sides as configured like this. No negotiation takes place.

## PAgP

PAgP is a Cisco Proprietary protocol. Up to 8 ports of the same type can be configured for the same EtherChannel.

```
Sw(config)# channel-group PO-ID mode {auto|desirable} [non-silent]
```

* **auto** – enables PAgP passively – It will respond to PAgP packets, but will not start a negotiation.
* **desirable** – enables PAgP actively – it responds and starts a PAgP negotiation.
* *silent* – by default, PAgP assumes the silent mode, when it considers that the other end doesn’t use PAgP so it will bring the portchannel Up, if the other end doesn’t respond to PAgP packets.
* **non-silent** – If set, then PAgP will wait for a response from the other before bringing up the PortChannel.

## LACP

LACP is a IEEE protocol. Up to 16 ports of the same type can be configured in an EtherChannel, but only 8 will be active, while the other 8 will be in standby. The software will determine which ports to use based on the lowest priority, which consists of:

1. LACP System Priority
2. System ID (Switch MAC)
3. LACP Port Priority
4. Port Number

The side with the lowest System Priority will decide what ports to use. To modify and verify the System Priority, use:

```
Sw(config)# lacp system-priority PRI
!Default: 32768
Sw# show lacp sys-id
```

To modify and veridy the LACP Port Priority, use:

```
Sw(config-if)# lacp port-priority PRI
!default: 32768
Sw# show lacp [PO#] internal
```

To set up an EtherChannel using LACP, use:

```
Sw(config)# channel-group PO-ID mode {active|passive}
```

* **passive** – enables LACP passively – It will respond to LACP packets, but will not start a negotiation.
* **active** – enables LACP actively – it responds and starts a LACP negotiation.

## Load Balancing

In order to decide on which link should traffic be forwarded out on a channel group, the switch uses a load balancing algorithm which takes the source or destination MAC or IP address of each packet and computes a hash value which is mapped to the port it should be forwarded on. As long as the input data is the same, the same port will be chosen for output.\
To set the load-balance method, use:

```
Sw(config)# port-channel load-balance {dst-ip|dst-mac|src-dst-ip|src-dst-mac|src-ip|src-mac}
!Default: src-mac
```

To verify, use:

```
Sw# show etherchannel load-balance
```

The hash is fixed length and most platforms use a 3 bit size. This creates some unbalance in utilization as the ports are mapped to the possible hash values:

| Ports in EtherChannel | Hash assignment        | Ratio           |
| --------------------- | ---------------------- | --------------- |
| 2                     | 1\|2\|1\|2\|1\|2\|1\|2 | 4:4             |
| 3                     | 1\|2\|3\|1\|2\|3\|1\|2 | 3:3:2           |
| 4                     | 1\|2\|3\|4\|1\|2\|3\|4 | 2:2:2:2         |
| 5                     | 1\|2\|3\|4\|5\|1\|2\|3 | 2:2:2:1:1       |
| 6                     | 1\|2\|3\|4\|5\|6\|1\|2 | 2:2:1:1:1:1     |
| 7                     | 1\|2\|3\|4\|5\|6\|7\|1 | 2:1:1:1:1:1:1   |
| 8                     | 1\|2\|3\|4\|5\|6\|7\|8 | 1:1:1:1:1:1:1:1 |

Other platforms use 8 bits for hashing which smooth out some of the unbalance

## Monitor

```
Sw# show ethernchannel [PO-ID] {summary|detail|...}
Sw# show pagp ...
Sw# show lacp ...
```

### EtherChannel Guard

This feature can detect misconfigurations between connected devices. Normally it is configured by default and it will put the physical interfaces in *err-disabled* state.

To check if Etherchannel Guard is enabled, use:

```
Sw# show spanning-tree summary
Switch is in pvst mode
Root bridge for: VLAN0001
Extended system ID           is enabled
Portfast Default             is disabled
PortFast BPDU Guard Default  is disabled
Portfast BPDU Filter Default is disabled
Loopguard Default            is disabled
EtherChannel misconfig guard is enabled
...
```

To enable the feature, use the global config command:

```
Sw(config)# spanning-tree etherchannel guard misconfig
```


# Layer 2 WAN Protocols


# HDLC


# HDLC 101

## Basic Configuration

The default encapsulation on serial interfaces on a Cisco Router is HDLC.\
Even though HDLC exists as an open standard, Cisco routers use a modified version of the original standard that includes a new field in the header used for identification of the upper layer protocol that is encapsulated. Many other vendors can interoperate with Cisco Routers even when using this version of the protocol, but technically it is considered a non-standard implementation. If the requirements are to implement a standard encapsulation on a Serial interface, a better option would be PPP.

Since HDLC is the default encapsulation, configuration is probably the easiest ever:

```
! On R1:
R1(config)# interface Serial0/0
R1(config-if)# ip address 99.0.0.1 255.255.255.0
R1(config-if)# no shut
! On R2:
R1(config)# interface Serial0/0
R1(config-if)# ip address 99.0.0.2 255.255.255.0
R1(config-if)# no shut
! Test
R1# ping 99.0.0.2
!!!!!
```

Normally, when connecting two routers on a serial link, one of them must be the DTE, the other must be the DCE.\
The DCE router must set a clock rate for the interface using:

```
R1(config-if)# clock rate CLOCK-RATE
```

In newer versions of the IOS, this is not needed anymore because by default, every serial interface has a clock rate set, but only the DCE end will use it.

How do we know which is the DCE end? The easiest method is to look at the command:

```
R1# show controllers serial0/0 | i DCE|DTE
cable type : V.11 (X.21) DCE cable, received clockrate 2015232
```

## Keepalives

Cisco HDLC uses keepalives to monitor the link state. By default, keepalives are sent and expected every 10 seconds. Three missed keepalives will move the interface protocol to a “down” status.

```
R1(config-if)#keepalive ?
< 0-32767> Keepalive period (default 10 seconds)
< cr>
```

Hitting enter will set the default value of 10 seconds. To disable keepalives completly use:

```
R1(config-if)# no keepalive
```

To verify keepalive value use:

```
R1# show interface serial0/0 | i alive
Keepalive set (10 sec)
```

## Compression

HDLC can be configured to use software compression using the Stacker(LZS) algorithm. Compression must be enabled on both ends of the link:

```
!On R1
R1(config-if)# compress stack
!On R2
R2(config-if)# compress stack
```

Compression can be verified using:

```
R1# show compress [detail]
Serial0/0
Software compression enabled
uncompressed bytes xmt/rcv 101850/100816
compressed bytes xmt/rcv 43041/42357
Compressed bytes sent: 43041 bytes 3 Kbits/sec ratio: 2.366
Compressed bytes recv: 42357 bytes 3 Kbits/sec ratio: 2.380
1 min avg ratio xmt/rcv 2.152/2.139
5 min avg ratio xmt/rcv 2.150/2.137
10 min avg ratio xmt/rcv 2.150/2.137
no bufs xmt 0 no bufs rcv 0
resyncs 0
Additional Stac Stats:
Transmit bytes: Uncompressed = 220 Compressed = 43041
Received bytes: Compressed = 46381 Uncompressed = 0
```

Notice the compression ratio. One byte was sent after compression for each 2.366 bytes of uncompressed data. In the other direction, 1 byte of data was received for each uncompressed 2.380 bytes of date. These are of course average values. Beware that compression can affect system performance!


# PPP


# PPP 101

PPP is an open-standard, media-independent Layer 2 protocol that can provide such features as authentication, multilink, compression, reliability.

PPP uses 2 negotiation phases. First, LCP (Link Control Protocol) is used to set up the link and to negotiate authentication, compression mechanisms, available MTU size, and so on, and then there’s a NCP phase (Network Control Protocol) specific for each network layer protocol. There’s IPCP – for IPv4, IPv6CP – for IPv6 or CDPCP – for CDP (CDP is a actually a Layer 2 protocol that works over PPP)

For example, after LCP phase is over, the IPCP protocol will be used to exchange IP related information, like getting the IP address of the interface.

## Physical interface vs Dialer Interface

### Physical interface

For back to back connections between 2 routers, the easiest way is to use PPP encapsulation on the physical interfaces connecting them:

```
R(config-if)# encapsulation ppp
R(config-if)# no shut
```

### Dialer Interface

In other environments, you can define the server’s physical interface to use PPP encapsulation as above, and use a Dialer Interface on the client.

```
R(config)# interface DIALER-INT
R(config-if)# encapsulation ppp
! By default, dialer interfaces use HDLC encapsulation
R(config-if)# dialer-pool POOL-ID
```

The connection between the Dialer interface and physical interfaces is done via a Dialer Pool

![PPP Dialer](/files/syGxyBFcPBxAfWPppRfe)

To configure the physical interface to be a member of a Dialer Pool, use:

```
R(config)# interface SERIAL-INT
R(config)# encapsulation ppp
R(config)# dialer in-line
R(config-if)# dialer pool-memeber POOL-ID
```

From now on, all PPP configuration is done on the Dialer interface, which acts as a PPP Profile that is applied on the physical interfaces.

A Dial interfaces is activated whenever interesting traffic is sent over it. To define the interesting traffic you will have to configure the Dial interface to be part of a Dial Group.

```
R(config-if)# dialer-group DIALER-LIST
```

Then, the Dial Group will reference a Dial List, which will define interesting traffic. This configuration is done globally:

```
R(config)# dialer-list DIALER-LIST protocol PROTOCOL {list ACL|permit|deny}
! PROTOCOL - ip, ipv6, bridge, etc
! permit - will define all PROTOCOL traffic as interesting traffic
! deny - will remove all PROTOCOL traffic from the interesting traffic
! list ACL - will define interesting traffic based on ACL
```

Whenever there is interesting traffic to be sent over the Dial interface, it will generate a Dial Call.\
You can configure the Dial interface to be always up if you use:

```
R(config-if)# dialer persistent
! To keep the interface up/up even if there is no interesting traffic
! This command will auto-generate:
R(config-if)# dialer idle-timeout 0
```

You can verify the status of a Dialer interface using:

```
R# show dialer
```

## Keepalives

PPP uses keepalives to monitor the link state. By default, keepalives are sent and expected every 10 seconds. Unlike most protocols, **five** missed keepalives will move the interface protocol to a “down” status.

```
R(config-if)#keepalive ?
&lt; 0-32767&gt; Keepalive period (default 10 seconds)
&lt; cr&gt;
```

Hitting enter will set the default value of 10 seconds. To disable keepalives completly use:

```
R1(config-if)# no keepalive
```

To verify keepalives, use:

```
R1# show interface serial0/0 | i alive
Keepalive set (10 sec)
```

## IP address Pooling

### Client Config

A PPP client can manually set its IP address with

```
R(config-if)# ip address IP-ADDR NETMASK [secondary]
```

But it can also use IPCP to discover the IP address it should use. To enable the use of IPCP for address assginement, use:

```
R(config-if)# ip address negotiated [previous]
```

If the **previous** keyword is used, then the client will attempt to get the previously assigned address from the server. The server can ignore this request if it uses:

```
R(config-if)# peer ip address forced
```

Addresses assigned using IPCP are always /32 addresses. Additional information may be requested by the client if it is explicitly configured:

```
R(config-if)# ppp ipcp dns request
R(config-if)# ppp ipcp mask request
! While this is not used by PPP, it may be used by other functions like DHCP ODAP
R(config-if)# ppp ipcp wins request
```

In order to reply to these requests, the server must be explicitly configured to send them. See next section.

### Server Config

When a PPP client requests its IP Address from the PPP Dial-in server, the latter will send by default an IP address from a local pool. This default mechanism can be overridden globally or on the interface,\
Globally:

```
R(config)# ip address-pool {local|dhcp-pool|dhcp-proxy-client}
! local - Uses the default local pool (Default)
! dhcp-pool - Uses a local DHCP pool
! dhcp-proxy-client - Proxies the request to a DHCP server 
```

Per interface:

```
R(config-if)# peer default ip address {CLIENT-IP | pool [IP-POOL]| dhcp-pool [DHCP-POOL] |dhcp}
! CLIENT-IP: assigns the configured address to clients on this interface
! pool [IP-POOL]: Sets a local-pool as default mechanism on this interface.
! dhcp-pool [DHCP-POOL]: Uses a local DHCP Pool as the default mechanism on this interface
! dhcp: Acts a a DHCP Proxy Client and forwards requests to a remote DHCP Server
```

#### **Local Address Pools**

To define a local IP-POOL, then use:

```
R(config)# ip local pool {IP-POOL|default} START-IP [END-IP] ...
```

#### **Local DHCP Pools**

To define a DHCP-POOL and enter DHCP Pool Configuration Mode, use:

```
R(config)# ip dhcp pool DHCP-POOL
```

#### **DHCP Proxy**

For DHCP Proxy client, configure the remote DHCP Server, using:

```
 R(config)# ip dhcp-server IP-ADDR
! Optionally, allow DNS and Netbios discovery:
R(config)# ip dhcp-client network-discovery informs INFORMS discovers DISCOVERS period SEC
```

There is a problem when using the DHCP client. The router will forward the request sourced from the PPP interface. The DHCP server will send the reply unicast to this address. But since the negotiation is not completed until the other side receives an address (or renounces after too many fails – which will be too late anyway), the interface will have an *up/down* status. This means that it won’t be advertised by routing protocols, so probably the DHCP server doesn’t have a route back to it. This can be resolved by using static routes on the path from the server to our router, or in a more dynamic way, by using a loopback address and setting the address on the PPP interface as unnumbered from that loopback. This will let the prefix to be advertised by routing protocols.

#### **Responding to IPCP requests**

When an IPCP client requests additional information, the router has to explicitly define them with:

```
R(config-if)# ppp ipcp dns SERVER1 [SERVER2...]
R(config-if)# ppp ipcp netmask MASK
R(config-if)# ppp ipcp wins SERVER1 [SERVER2...]
```

## PPP host route

When a PPP link is up, it will add a /32 route to that points to the IP address on the other end of the PPP link. This makes it possible for hosts with IP addresses in different subnets to be directly connected over a PPP link.

```
R2#sh ip route
Gateway of last resort is not set

     1.0.0.0/32 is subnetted, 1 subnets
C       1.1.1.1 is directly connected, Serial1/1
     2.0.0.0/24 is subnetted, 1 subnets
C       2.2.2.0 is directly connected, Serial1/1
```

This feature can be disabled using the command:

```
R2(config-if)# no peer neighbor-route
```

## Authentication

PPP supports multiple authentication protocols, like [PAP](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-authentication-pap), [CHAP](/layer-2-technologies/layer-2-wan-protocols/ppp/ppp-authentication-chap), MS-CHAP, MS-CHAPv2 and EAP.

## Compression

PPP can be configured to use a software compression algortithm. Compared to HDLC, PPP supports more compression algorithms and can be configured with different algorithms on each end. PPP LCP phase will take care of the negotiation.

To see the available compression algorithms use:

```
R(config-if)#compress ?
  lzs        lzs compression type
  mppc       MPPC compression type
  predictor  predictor compression type
  stac       stac compression algorithm
  < cr>
```

To see compression statistics use:

```
R1#sh compress details
Serial1/3
Software compression enabled
uncompressed bytes xmt/rcv 15627/15729
compressed bytes xmt/rcv 3868/4258
Compressed bytes sent: 3868 bytes 0 Kbits/sec ratio: 4.040
Compressed bytes recv: 4258 bytes 0 Kbits/sec ratio: 3.693
1 min avg ratio xmt/rcv 0.130/0.131
5 min avg ratio xmt/rcv 0.061/0.062
10 min avg ratio xmt/rcv 0.061/0.062
no bufs xmt 0 no bufs rcv 0
resyncs 0
Compression Protocol used: PPP Stac LZS
Check Mode = 3
Decompression Protocol used: Predictor 1
```

Notice the compression ratio and that different protocols are used for compression (TX) and decompression (RX)


# PPP Authentication - PAP

PAP is a simple but not very secure authentication protocol. It sends the username and password information in clear text.

## One-way authentication

On the router that asks for authentication:

```
R1(config)# username USER password PASS
R1(config)# interface serial0/0
R1(config-if)# encapsulation ppp
R1(config-if)# ip address IP-ADDR1 MASK1
R1(config-if)# ppp authentication pap
R1(config-if)# no shut
```

On the router that is authenticating:

```
R2(config)# interface serial0/0
R2(config-if)# encapsulation ppp
R2(config-if)# ip address IP-ADDR2 MASK2
R2(config-if)# ppp pap sent-username USER password PASS
R2(config-if)# no shut
```

This is clearly a one-way authentication. Only R2 authenticates to R1.

## Two-way authentication

If two-way authentication is required, a configuration like the following should be used:

```
! On R1:
R1(config)# username USER1 password PASS1
R1(config)# interface serial0/0
R1(config-if)# encapsulation ppp
R1(config-if)# ip address IP-ADDR1 MASK1
R1(config-if)# ppp authentication pap
R1(config-if)# ppp pap sent-username USER2 password PASS2
R1(config-if)# no shut
! On R2:
R1(config)# username USER2 password PASS2
R2(config)# interface serial0/0
R2(config-if)# encapsulation ppp
R2(config-if)# ip address IP-ADDR2 MASK2
R2(config-if)# ppp authentication pap
R2(config-if)# ppp pap sent-username USER1 password PASS1
R2(config-if)# no shut
```

## Other Settings

By default a router will always respond to an authentication request even when no username is configured to be sent, but authentication will fail. Use the following command to refuse authentication requests:

```
R(config-if)#ppp pap refuse
```

### Debugging

When debugging, the best commands to use are:

```
R# debug ppp negotiation
R# debug ppp authentication
```

See this Cisco article about [debugging PPP neogtiation output](https://www.cisco.com/en/US/tech/tk713/tk507/technologies_tech_note09186a00800ae945.shtml)


# PPP Authentication – CHAP

## CHAP algorithm

CHAP is more secure than PAP because even though it sends usernames in clear text, it uses an MD5 hash instead of a clear text password for authentication.

When a router is set to ask for authentication, it will send a CHAP Challenge to the other router which contains an **ID**, a **random number** and a **username**. By default, the username used is the hostname of the router. For each Challenge, the router will store in memory the random number used.

When the router that shoud authenticate receives the Challenge, it must search in it’s database a password for the supplied username. If it doesn’t find one, the authentication will fail. If if finds one, the router will use the **ID**, **the random number** and the **password** to generate an **MD5 hash**. It will send this MD5 hash in a Response packet, together with a **username** to the router that issued the authentication challenge.The username is by default the hostname of the router that responds

Upon receiving the Response, the first router uses the ID to determine what Challenge the Response belongs to and retrives the random number used in the Challenge it sent. It then searches for a password coresponding to the username received from the other router. If no match is found, authentication fails. If a match is found, then the router uses the **ID**, **the random number** it used in the Challenge and the **password** it found to generate an **MD5 hash**. If the MD5 hash it generates matches the one that it received, then it (statisticly) means both routers used the same password and authentication succeeds.

## One-way authentication

The configuration for one-way authentication should look like this:

```
!On R1
R1(config)# interface serial0
R1(config-if)# ip address IP-ADDR1 MASK1
R1(config-if)# encapsulation ppp
R1(config-if)# ppp authentication chap
R1(config-if)# exit
R1(config)# username R2 password PASS
!On R2
R2(config)# interface serial0
R2(config-if)# ip address IP-ADDR2 MASK2
R2(config-if)# encapsualation ppp
R2(config-if)# no shut
R2(config-if)# exit
R2(config)# username R1 password PASS
```

Even though this might look like a two way authentication, it isn’t. For two-way authentication both routers should send a Challenge.

## Two-way authentication

```
!On R1
R1(config)# interface serial0
R1(config-if)# ip address IP-ADDR1 MASK1
R1(config-if)# encapsulation ppp
R1(config-if)# ppp authentication chap
R1(config-if)# exit
R1(config)# username R2 password PASS
!On R2
R2(config)# interface serial0
R2(config-if)# ip address IP-ADDR2 MASK2
R2(config-if)# encapsualation ppp
R2(config-if)# ppp authentication chap
R2(config-if)# no shut
R2(config-if)# exit
R2(config)# username R1 password PASS
```

As you can see, when using two way authentication you cannot use different passwords, as was possible with PAP.

## More settings

To use a different username than the configured hostname, you can use:

```
R(config)# username USER2 password PASS2
R(config)# interface serial0
R(config-if)# ppp chap hostname USER2
```

and to use a default password for unknown usernames, you can use:

```
R(config)# interface serial0
R(config-if)# ppp chap password PASS
```

By default a router will always respond to an authentication request with its hostname as a username, even if there is no password configured. Of course, authentication will fail. Use the following command to refuse authentication requests:

```
R(config-if)#ppp chap refuse
```

## Debugging

When debugging, the best commands to use are:

```
R# debug ppp negotiation
R# debug ppp authentication
```

See this Cisco article about [debugging PPP neogtiation output](https://www.cisco.com/en/US/tech/tk713/tk507/technologies_tech_note09186a00800ae945.shtml)


# PPP Authentication – EAP

## One way authentication

### EAP Server (Authenticator)

```
! Set encapsulation to PPP
R(config-if)# encapsulation ppp
! Enable PPP authentication
R(config-if)# ppp authentication eap
```

PPP will require the use of a Radius. However, you can use the local usernames if you configure:

```
R(config-if)# ppp eap local
! Use local username and passwords to perform authenticaion
```

Remember to configure the username and password that the client will use:

```
R(config)# username USER password PASS
```

### EAP Client (Authenticating)

```
R(config-if)# encapsulation ppp
! Set the password used for EAP
R(config-if)# ppp eap password PASS
! By default, the hostname is used as the username. You can change it with:
R(config-if)# ppp eap identity USER
```

## Two way authentication

On R1:

```
R1(config)# username USER2 password PASS2
R1(config)# interface INTERFACE
R1(config-if)# encapsulation ppp
R1(config-if)# ppp authentication eap
R1(config-if)# ppp eap local
R1(config-if)# ppp eap password PASS1
R1(config-if)# ppp eap identity USER1
```

On R2:

```
R1(config)# username USER1 password PASS1
R2(config)# interface INTERFACE
R2(config-if)# encapsulation ppp
R2(config-if)# ppp authentication eap
R2(config-if)# ppp eap local
R2(config-if)# ppp eap password PASS2
R2(config-if)# ppp eap identity USER2
```


# PPP Multilink

## PPP Mulitlink

One feature that PPP can offer is the posibility to bundle multiple physical links into one logical link, much like an EtherChannel interface available on the Ethernet Switches.\
Considering two routers connected back-to-back over more than one serial links. Each serial interface can be configured for PPP encapsulation and assigned to a Multilink Group:

```
R(config-if)# encapsulation ppp
R(config-if)# ppp multilink group GROUP
```

Then, you can use the virtual Multilink interface to assign an IP Address:

```
R(config)# interface Multilink GROUP
R(config-if)# ppp multilink
R(config-if)# ip address IP-ADDR NETMASK
```

By default the peer neighbor-route feature is eanbled, and we will see /32 routes in the routing table.

The link will be operational as long as one link is active. This behaviour can be changed using the command:

```
R1(config-if)#ppp multilink links minimum MIN [mandatory]
```

*MIN* specifies the minimum number of links needed to be up before the multilink interface goes up. When the **mandatory** keyowrd is used, if the number of active links goes below the number of minimum links, the multilink interface goes down. If **mandatory** is not used, the minimum links value is checked only when the interface first comes up. If the number of active links drops below minimum while the interface is up, the multilink will still remain up as long as one link is operational.

You can check the status of the multilink using:

```
R1#sh ppp multilink
Multilink1, bundle name is R2
Endpoint discriminator is R2
Bundle up for 00:14:20, total bandwidth 3088, load 1/255
Receive buffer limit 24000 bytes, frag timeout 1000 ms
0/0 fragments/bytes in reassembly list
30 lost fragments, 53 reordered
16/918 discarded fragments/bytes, 0 lost received
0x89 received sequence, 0x2E sent sequence
Member links: 2 active, 0 inactive (max not set, min not set)
Se1/2, since 00:14:20
Se1/0, since 00:12:54
No inactive multilink interfaces
```

## PPP Multilink over Frame Relay

PPP Multilink over Frame Relay is configured similarly. The steps to enable it are:\
1\. Configure the Serial interfaces for Frame Relay encapsulation

```
R(config)# interface INTERFACE
R(config-if)# encapsulation frame-relay
```

2\. Configure the Serial interfaces for PPP:

```
R(config-if)# frame-relay interface-dlci DLCI ppp VIRTUAL-TEMPLATE-INTERFACE
```

3\. Configure the Virtual Template to be part of the Multilink Group

```
R(config)# interface VIRTUAL-TEMPLATE-INTERFACE
R(config-if)# ppp multilink
R(config-if)# ppp multilink group GROUP
```

4\. Configure the Multilink interface

```
R(config)# interface MULTILINK-INTERFACE
R(config-if)# ppp multilink
R(config-if)# ppp multilink group GROUP
R(config-if)# ip address...
```


# PPPoFR – PPP over Frame Relay

## PPPoFR

PPP can be used encapsulated in Frame Relay to offer all PPP features over an existing Frame Relay network. The first feature that comes to mind is probably authentication, but PPPoFR can also be used to offer advantages such as Multilink PPP over Frame Relay or PPP compression. Also, PPP can help in situations where point-to-point subinterfaces or inverse ARP cannot be used. PPPoFR will use IPCP instead of Frame Relay’s Inverse ARP to determine how to send an IP packet over the network. It also bypasses issues related to Split Horizons.\
In order to configure PPPoFR we will need to use a Virtual Template interface and a Virtual Access interface. The Virtual Template interface acts as a PPP interface and all configuration is done on it, but it will always show as down/down. The Virtual Access interface will get its configuration from the Virtual Template and will be the interface that will be up/up if everything works out.

To enable PPPoFR, follow these steps:\
1\. Enable Frame Relay encapsulation

```
R(config)# interface SERIAL-INT
R(config-if)# encapsulation frame-relay
```

2\. Enable PPP on the frame-relay interface:

```
R(config-if)# frame-relay interface-dlci DLCI ppp VIRTUAL-TEMPLATE-INT
```

3\. Then, configure the VIRTUAL-TEMPLATE interface with an IP address and with any other PPP options:

```
R(config)# interface VIRTUAL-TEMPLATE-IN
R(config-if)# ip address ...
R(config-if)# ppp ...
```

To verify, you can look at the routing table

```
R3#sh ip route
Gateway of last resort is not set

     2.0.0.0/32 is subnetted, 1 subnets
C       2.2.2.2 is directly connected, Virtual-Access1
     3.0.0.0/32 is subnetted, 1 subnets
C       3.3.3.3 is directly connected, Loopback0
```

As you can see, there’s the classic /32 ip route inserted by PPP, but it shows up connected on the Virtual-Access1 interface, not on Virtual-Template1. Also, look at the command below to see that the Virtual-Template interface is down/down while the Virtual-Access interface is up/up.

```
R3#sh ip int brief | i Virtual
Virtual-Access1            3.3.3.3         YES TFTP   up                    up
Virtual-Template1          3.3.3.3         YES TFTP   down                  down
Virtual-Access2            unassigned      YES unset  down                  down
```

## Ping yourself

To ping yourself in Frame Relay you must have a frame-relay map that points to your own IP. With PPPoFR it gets even more complicated, because the configuration uses a Virtual Acces interface that copies its configuration from the Virtual Template interface. Due to the fact that the Virtual Template interface is always down/down, ping to yourself will fail. The solution is to use ip unnumbered from a Loopback interface or make the virtual-template part of a PPP multilink interface.


# PPPoE – PPP over Ethernet

PPP can be encapsulated over Ethernet and it can be used to offer authentication over Ethernet connections. It cannot be used for Multilinking Ethernet connections. When using PPPoE we need a server and a client.

## Server Side

Configuring the server makes use of the Virtual Template/Virtual Access interfaces. We configure the Virtual Template interface, but the Virtual Access interface will be up/up in the process. Its configuration will be taken from the Vitual Template interface.\
First define the Virtual Template that will hold the PPP configuration:

```
R1(config)# int VIRTUAL-TEMPLATE-INT
R1(config-if)# ip address IP-ADDR NETMASK
! define IP assignment for clients:
R1(config-if)# peer default ip address {CLIENT-IP | pool [IP-POOL]| dhcp-pool [DHCP-POOL] |dhcp}
```

See [PPP 101](https://nyquist.eu/ppp-101/#32_Server_Config) for details about client IP assignment.\
Then, configure the Broadband Access Group that points to the Virtual-Template interface:

```
R1(config)# bba-group pppoe {GROUP-NAME|default}
R1(config-bba-group)# VIRTUAL-TEMPLATE-INT
```

The last step is to enable PPPoE on the physical interface and assign it to the BBA-Group:

```
R1(config)# interface FAST-ETHERNET-INT
R1(config-if)# pppoe enable group {GROUP|default}
```

## Client Side

On the client side, we have to make use of Dialer Interfaces. Dialer Interfaces will look for a physical interface available in their Dialer Pool before initiating connections

![PPPoE Dialer](/files/syGxyBFcPBxAfWPppRfe)

```
R2(config)# interface DIALER1
! By default Dialer interfaces have encapsulation set to HDLC.
R2(config-if)# encapsulation ppp
!Get the IP address from the server
R2(config-if)# ip address {negotiated|IP-ADDR NETMASK}
R2(config-if)# dialer-pool POOL-ID
! What pool shoud I use?
```

Physical interfaces are assigned to Dialer Pools:

```
R2(config)# interface FASTETHERNET0
! Link the Ethernet Interface to the Dialer interface via the dial-pool.
R2(config-if)# pppoe-client dial-pool-number POOL-ID
! What pool am I a member of?
R2(config-if)# no shut
R2(config-if)# end
```

Additional information about Dialer Interfaces, can be found [here](https://nyquist.eu/ppp-101/#12_Dialer_Interface).

After a little while we will see the assigned IP address on the interface and the PPP /32 ip route:

```
R2#sh ip route
Gateway of last resort is not set

     1.0.0.0/32 is subnetted, 1 subnets
C       1.1.1.1 is directly connected, Dialer1
     99.0.0.0/32 is subnetted, 1 subnets
C       99.0.0.10 is directly connected, Dialer1

R2#sh ip int brie
Interface                  IP-Address      OK? Method Status                Protocol
FastEthernet0/0            unassigned      YES unset  up                    up
FastEthernet0/1            unassigned      YES unset  administratively down down
Virtual-Access1            unassigned      YES unset  up                    up
Dialer1                    99.0.0.10       YES IPCP   up                    up
R2#ping 1.1.1.1

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 1.1.1.1, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 12/24/52 ms
R2#
```

If we want to enable authentication, just configure the PPP interfaces on each device, that is the Virtual-Template on the Server and the Dialer interface on the Client.


# Frame Relay


# Frame Relay 101

Frame Relay is a Layer 2 WAN technology that uses the concept of virtual circuits to transport data. The Frame Relay network is packet switched, but packets towards a destination share a common label, so they follow a strict path according to the label, giving the impression of an end-to-end circuit to the upper layer protocols.

## Understand what device you are configuring

Usually, devices in a Frame Relay network have 2 main roles: The customer **Frame Relay Router** and the provider **Frame Relay Switch**. I like to call them like this because it makes the difference more clear between the 2 roles. One difference between a switch and a router is that, typically, a router makes forwarding decisions based on the Layer 3 address, and a switch based on the Layer 2 address. In Frame Relay, the customer router will receive or create a Layer 3 packet. It will look at its destination address and based on the routing table it will have to forward it over a Frame Relay interface. To do this, it will have to encapsulate the Layer 3 packet into a Layer 2 “frame” before sending it out the Frame Relay interface. The next device on the link is the provider Frame Relay Switch, which will forward the frame based only on the Layer 2 information, without looking at Layer 3. This is why the devices in the provider’s Frame Relay cloud are called Frame Relay Switches.

It is important to understand what device you are configuring because the configuration differs on a Frame Relay Router from the one on a Frame Relay Switch.

Most of the time you will see that the Frame Relay Router is called the DTE device, and the Frame Relay Switch is called the DCE device.

In Telecommunications, DTE is a Data Terminal Equipment and is the last device in a communication line, while DCE is the Data Communication Equipment and is the device that the DTE connects to, in order to access the communication path. This is why the interface on the Frame Relay Router is defined as DTE interface, while the interface on the Frame Relay Switch is defined as DCE interface. What about interfaces between Frame Relay Switches? They are called NNI – Network to Network Interface.

Let me point out that this denomination is different from the DTE and DCE ends of a serial connection, because even though the Frame Relay DCE device is usually the DCE end of the link and the DTE device is usually the DTE link, the connection can still work if the Frame Relay DTE device is the DCE end of the link and the Frame Relay DCE device is the DTE end of the link

![Frame Relay 101](/files/Pwfw6KogYjYhiaRrL9Hc)

## DLCI (Data Link Connection Identifier)

As I said earlier, the Frame Relay network offers a virtual circuit to the upper layer protocols, by encapsulating the Layer 3 information in the Layer 2 frames, and by marking each frame with its appropriate Data Link Connection Identifier. The name says it all. The DLCI will identify the virtual circuit that each frame belongs to. One thing that must be remembered here is that the DLCI identifies the circuit on a link by link basis. This means that the the same virtual circuit will be identified by different DLCI values on each link from the source to the destination.

In the example below, R1 uses 2 different circuits to reach R2 and R3 respectively. R1 will encapsulate traffic for R2 in frames with DLCI set to 102 and will send them towards FRS1. FRS1 knows it must take frames with DLCI 102 and send them towards FRS0 with DLCI set to 912. When FRS0 receives frames with DLCI 912 will switch them over to FRS2 and will change the DLCI to 452. As the last device in the Frame Relay Cloud, FRS2 will forward the frames to R2 but it will also replace the DLCI value with 201. The same process happens for traffic from R1 to R3, but this time the DLCIs used are 103, 913, 743 and 301

![Frame Relay DLCI](/files/Pwfw6KogYjYhiaRrL9Hc)

## Encapsulation types – CISCO and IETF

To complicate thinks a bit, there are different type of encapsulations that can be used. Cisco routers default to ***cisco*** encapsulation, but there is also an ***ietf*** encapsulation available, that should be used when connecting to non-Cisco devices. Luckily, many other devices support *cisco* encapsulation, but also, Cisco devices work even with a mismatched encapsulation. They will send packets with the defined encapsulation but will accept packets with either encapsulation. So in real life there is no need to worry about this, but it is recommended to match the encapsulation even for the often underrated reason of readability.

```
R(config)# interface serial1/0
R(config-if)#encapsulation frame-relay ?
 MFR Multilink Frame Relay bundle interface
 ietf Use RFC1490/RFC2427 encapsulation
 
```

Even if it is not very clear, we should understand that hitting enter will set the encapsulation to *cisco*, while the other options are *ietf* and [*MFR*](https://nyquist.eu/multilink-frame-relay/).

To see the encapsulation type used on an interface use:

```
R1#show interface s1/0 | i Encapsulation
 Encapsulation FRAME-RELAY, crc 16, loopback not set
```

Another important topic here is that the frame-relay encapsulation must match **end-to-end**, that is between the DTE devices, and not between the DTE and the DCE. There is something else that should match between the DTE and the DCE and that is the LMI Type.

You will see later that encapsulation can be different on each virtual circuit, but by default, each virtual circuit will inherit the settings on the interface. To see how encapsulation actually looks like, you should read this article about [Frame Relay Encapsulation](https://nyquist.eu/frame-relay-encapsulations-ietf-vs-cisco/)

## LMI types – CISCO, ANSI, Q933A

Frame Relay started in the mid 1980’s when both [ITU-T](https://en.wikipedia.org/wiki/ITU-T) (known at the time as CCITT) and [ANSI](https://en.wikipedia.org/wiki/American_National_Standards_Institute) where trying to standardize the technology, but it took off at the beginning of the 1990’s when “The Gang of Four” – Cisco, StrataCom, Northern Telecom and DEC – created a consortium that would focus on accelerating the introduction and interoperability of Frame Relay products. The most important development they came up with, was the LMI – Local Management Interface, which is a set of extensions complementary to the  existing Frame Relay standards that add several new features to it, like Virtual Circuit Status Messages, Multicasting or Global Addressing. Support for Multicasting and Global Addressing extensions is optional, but VC Status Messages is expected to be implemented by most vendors.

VC Status Messages offers a very useful service for Frame Relay, that gives an end-to-end status of the virtual circuit from one DTE to another. The status of a virtual circuit can be ACTIVE, INACTIVE or DELETED.

For a quick view of the status of each virtual circuit, I use:

```
R1#show frame-relay pvc | i DLCI
DLCI = 102, DLCI USAGE = SWITCHED, PVC STATUS = ACTIVE, INTERFACE = Serial1/0
DLCI = 103, DLCI USAGE = SWITCHED, PVC STATUS = ACTIVE, INTERFACE = Serial1/0
```

An **ACTIVE** status means everything is OK and the virtual circuit can be used to send data from one DTE to another.

An **INACTIVE** status means the virtual circuit is not completed, but the problem is not on the local link between the DTE and the DCE, but beyond the first DCE, in the provider network or at the other end of the connection.

A **DELETED** status means the DTE and the DCE do not have the same information regarding the virtual circuit. Most often, the virtual circuit on the DTE is configured with a different DLCI than on the DCE.

To set the LMI type, you use the command below to select either Cisco, ANSI Annex D or ITU Q933-A (Annex A)

```
R1(config-if)#frame-relay lmi-type ?
 cisco
 ansi
 q933a
```

but probably the best option would be to just let the default **LMI auto-sense** to work. This feature, available on Cisco routers, will discover the LMI type used by the DCE device and will use it. To enable auto-sense, if an lmi-type has been already set, use:

```
R1(config-if)#no frame-relay lmi-type
```

To see the LMI type that is used:

```
FRS1#show frame-relay lmi | i interface
 LMI Statistics for interface Serial1/0 (Frame Relay NNI) LMI TYPE = CISCO
 LMI Statistics for interface Serial1/1 (Frame Relay DCE) LMI TYPE = CISCO
 LMI Statistics for interface Serial1/2 (Frame Relay DCE) LMI TYPE = CISCO
```

You can also see what DLCI is used for LMI, with:

```
R3#sh int s1/1 | i LMI DLCI
  LMI DLCI 1023  LMI type is CISCO  frame relay DTE
```

| LMI Type | Listens on | Available DLCIs |
| -------- | ---------- | --------------- |
| Cisco    | 1023       | 16-1007         |
| ANSI     | 0          | 16-991          |
| ITU      | 0          | 16-991          |

LMI Status inquires are sent every 10 sec by default (show as type1 in debug frame LMI)

```
R(config-if)# keepalive SEC
! Default: 10
```

Full Status updates are sent every 6th inquiry (show as type0 in debug frame LMI)

```
R(config-if)# frame-relay lmi-n391dte COUNT
! Default: 6
```

DTE will report the status of each configured DLCI. The MTU size limits the number of DLCIs on a link. When using an MTU of 1500 Bytes a maximum only 296 DLCIs can be included in one status message. If the DCE device doesn’t receives 3 LMI Status messages it considers the link down.

## Back-to-Back Frame Relay

![Back-to-Back Frame Relay](/files/jgcwCG0Hky33vLP1GxJ0)

The Back-to-Back Frame Relay configuration is one where, God knows why, you would want to connect to routers on a serial link using Frame Relay encapsulation. First, we will have to enable frame-relay encapsulation and set a Layer 3 address

```
! On R3
R3(config-if)# encapsulation frame-relay
R3(config-if)# ip address 134.0.0.3 255.255.255.0
R3(config-if)# no shut
! On R4
R4(config-if)# encapsulation frame-relay
R4(config-if)# ip address 134.0.0.4 255.255.255.0
R3(config-if)# no shut
```

As we saw at the beginning, a Frame Relay router expects to be the DTE end of the link so it must connect to a DCE device. But this time, both devices are DTE devices. In order to make the link operational, one device must become the serial DCE device and this is done by setting clock rate on the interface. Use “?” to see all available clock rates.

```
R4(config-if)#clock rate CLOCKRATE
```

\* In newer version of the IOS, this is not needed anymore as all serial interfaces have a default clock rate, but only the DCE uses it.\
Each router still thinks it is a Frame Relay DTE, so by default it expects to communicate over LMI with the DCE device. Unless LMI is disabled, the link will not come up! So next thing to do is to disable LMI on both ends using:

```
R3(config-if)# no keepalive
R4(config-if)# no keepalive
```

One more thing must be done to achive connectivity. The routers must agree on the DLCI to use. This can be done by assigning a DLCI to the interface and using Inverse ARP, or by setting up a static mapping between IP and DLCI, but we’ll see how this works in the [next episode](https://nyquist.eu/frame-relay-101-part-2/).

## Frame Relay End-to-End Keepalives

End-to-end Keepalives are a different kind of keepalive, that test the end-to-end connectivity over Frame Relay, that is from one DTE to another. This way, the status of the PVC can be better monitored.\
To enable EEK, first you must define a map-class:

```
R(config)# map-class frame-relay MAP-CLASS
R(config-map-class)# frame-relay end-to-end keepalive mode {bidirectional|request|reply|passive-reply}
```

The EEK works by sending a Request on one side, and replying with a response to it on the other side. Each side can don both or any of these actions.

* **Bidirectional**: This end will both send requests and reply to requests. The other end must be bidirectional also in order to work. Used when the upstream path is different from the downstream path
* **Request**: This end will only send and wait for replies. The other end must be set to reply or passive-reply in order to work
* **Reply**: This end will only respond to requests. The other end must be set to Request in order to work
* **Passive-reply**: This end will reply but will not track the requests received. The other end must be set to Request in order to work

The keepalive configuration can be done with the following commands:

```
R(config-map-class)# frame-relay end-to-end keepalive error-threshold {recv|send} VALUE
! Default: 2 = number of errors in event-window that move the interface down
R(config-map-class)# frame-relay end-to-end keepalive event-window {recv|send} VALUE
! Default: 3 = how many events to monitor. Default: last 3
R(config-map-class)# frame-relay end-to-end keepalive success-events {recv|send} VALUE
! Defaut: 2 = number of errors in event-window that move the interface up
R(config-map-class)# frame-relay end-to-end keepalive timer {recv|send} VALUE
! Default: 10 sec
```

Finally, apply the map-class on the frame-relay interface, using:

```
R(config-if)# frame-relay class MAP-CLASS
```

To monitor, use:

```
R# sh frame-relay end-to-end keepalive [interface INTERFACE]
```


# Frame Relay 102

If you thought the [previous Frame Relay 101](/layer-2-technologies/layer-2-wan-protocols/frame-relay/frame-relay-101) would cover all basic aspects of Frame Relay, you were wrong. Here’s more info on Frame Relay and how it works, and believe me, we are looking just at the tip of the iceberg here :)

![Frame Relay – Hub and Spoke](/files/BiEYC9DEXWKyx4rxuEQO)

## Non-Broadcast. What difference does it make?

The omnipresent nowadays Ethernet networks rely on the broadcast properties of its medium. On Ethernet networks, there is an address that can be used to send frames to all hosts in a network. That is the FF:FF:FF:FF:FF:FF MAC Address. The first advantage of such an address is that a host that doesn’t know the Layer 2 address of a Layer 3 host can ask for it using this broadcast address. Everybody will receive the interrogation but only the holder of the Layer 3 address should respond to it. (I will not get into Proxy ARP or security issues here and I will assume that everybody plays nicely). When the first host receives the reply, it will have information needed to send the Layer 3 packet encapsulated in the Layer 2 Ethernet Frame. This looks simple enough and works like a charm but unfortunately, this doesn’t work on a Frame Relay network because there is no DLCI (Framer Relay Layer 2 address) that can be used to reach all other hosts. So how do we find out the Layer 2 address of a host?

### Static mapping

On Cisco routers you can statically define the DLCI to use for a certain destination. The command is applied on the interface and it looks like:

```
R(config-if)# frame-relay map PROTOCOL ADDRESS DLCI
! Example:
R(config-if)# frame-relay map ip 100.0.0.2 102
```

Why do I need to specify the protocol? That’s because Frame Relay can transport not just IPv4, but also IPv6 or CLNS. The address field is the Layer3 address of the destination according to the specified protocol. Here’s an example for an IPv6 mapping

```
R(config-if)# frame-relay map ipv6 2001::2 102
```

You can add more options after the DLCI. One of them is the encapsulation type. This will make the traffic on the specific virtual circuit to be encapsulated with either *cisco* or *ietf* format. The encapsulation set here can be different from the interface encapsulation. Actually, this is when you will most probably use it, when the traffic for a virtual circuit identified by a specific DLCI should not inherit the interface encapsulation.

```
R(config-if)# frame-relay map ip 100.0.0.2 102 {cisco|ietf}
```

Another option is the broadcast keyword. This enables support for pseudo-broadcast on Frame Relay. When the upper layer protocol knows it sends a broadcast packet, it will inform the Frame Relay process of this, and Frame Relay will send the frame on all the virtual circuits marked with the broadcast keyword.

```
R(config-if)# frame-relay map ip 100.0.0.2 102 broadcast
```

Now there’s a catch: In a Hub-and-Spoke topology, it is common for one spoke to use the path to the hub to reach the other spokes. This means that you end up using the same DLCI for each destination. The pseudo-broadcast feature will look at all the Layer 3 to DLCI mappings and when it will see the broadcast keyword it will send the frame out that virtual circuit. When having more mappings with the same DLCI, the packet will be sent several times on the same virtual circuit. Although this is not wrong, it is sub-optimal and in real life scenarios you should send the packet only once. How can we achive this? By marking only one mapping with the broadcast keyword. The other mappings involving the same DLCI will be set without the broadcast keyword.

### Inverse ARP aka Dynamic Mapping

Of course, we had to have an automatic mechanism for Layer 3 to Layer 2 mappings. Making just manual configurations could mean a lot of work in a large network. So we can’t use ARP like in Ethernet, but we can use another form of ARP, that is rightly named Inverse ARP. That is because in Frame Relay, when an interface is configured with an IP Address and a virtual circuit becomes active, the router will send an Inverse ARP packet asking who is on the other end of the link. It will also send information about it’s own address. The router on the other end will reply with it’s IP address and will also add the information it received to it’s frame-relay maps.

You can see a packet capture of the process in this [Frame Relay Inverse ARP Packet Capture](https://www.cloudshark.org/captures/87be3b4b6625). The capture lists 4 packets, but actually they are only one request and one response, but seen on both ends of the same virtual circuit. Notice that the DLCI changes from 201 in one frame to 102 on the other end. Also, you might be set off by the question “Who is 3091?” Well, if you look closely, 3091 is actually the HEX value of the entire [Q.922 address](https://nyquist.eu/frame-relay-encapsulations-ietf-vs-cisco/) in the first request. This value is carried within the ARP header over the frame relay cloud as the “Target Hardware Address”,  even though the DLCI changes from 201 to 102 in the Frame Relay header.

To see the Layer 3 to Layer 2 mappings use the command below. Mappings marked as dynamic are learned through Inverse ARP while those marked with static are manually configured. The example also shows different types of Layer 3 protocols used over the same DLCI. It si important to notice that the dynamically learned mappings support the pseudo-broadcast.

```
R1# show frame-relay map

Serial1/0 (up): ip 99.0.0.2 dlci 102(0x66,0x1860), dynamic,
broadcast,
IETF, status defined, active
Serial1/0 (up): ipv6 2002::6300:2 dlci 102(0x66,0x1860), static,
IETF, status defined, active
```

One more thing worth mentioning here is that Inverse ARP must find out what DLCIs it can use to find the other end’s IP address. In most configurations, where a router connects to the provider Frame Relay switch, Inverse ARP cannot work without LMI because LMI tells the router what virtual circuits are ACTIVE. However, when configuring Back-to-Back Frame Relay, that is 2 Frame Relay DTEs connected to eachother, LMI must be disabled but Inverse ARP will still work because the DLCI is assigned to the interface.

Inverse ARP requests are sent by default every 60 sec and the mappings are cached until they are cleared or the interface is reset. Also, remember that a static mapping will always overwrite a dynamic mapping!

## Frame Relay subinterfaces – point-to-point or multipoint

Every Frame Relay interface can have one or more subinterfaces. The important thing to remember is that subinterfaces can be point-to-point or multipoint. To define a subinterface, use:

```
R(config)# interface INTERFACE
R(config-if)# encapsulation frame-relay
R(config-if)# interface INTERFACE.SUBIF {point-to-point|multipoint}
```

What’s the difference?

A multipoint interface can have multiple virtual circuits assigned to it. Basically, there are multiple DLCIs that are used to transport traffic on a single interface. The physical interface and a multipoint subinterface enter in this category.

A point-to-point subinterface only uses one virtual circuit so only one DLCI is used to carry traffic. What is the advantage?

First of all, using separate point-to-point subinterfaces simplifies the design. One DLCI is assigned per sub-interface and traffic from one virtual circuit to another is routed as if there are separate physical interfaces. Plus, no need to worry about Split Horizons rules. In a multipoint interface updates that come in on one DLCI should be also advertised on the same interface on another DLCI. Layer 3 process doesn’t care about the DLCIs. It knows a routing update came on one interface, it will not send it back the same interface if split horizons rule apply (RIP and EIGRP).

The second advantage is that if there is only one DLCI per subinterface, then there is no need for Inverse ARP. Whatever we have to send, the interface will just encapsulate it with the assigned DLCI.\
On point-to-point subinterfaces you should also assign the DLCI:

```
R(config-subif)# frame-relay interface-dlci DLCI
```

You can only set one DLCI on point-to-point subinterfaces. In contrast, on multipoint interfaces or subinterfaces you can assign multiple DLCIs using the previous command several times. You will still have to map the Layer 3 to Layer 2 addresses in order to send the packet, otherwise encapsulation fails.\
When using LMI, all active DLCIs are associated to the physical interface. Using the above command you can assign DLCIs to other interfaces in order to help the Invers ARP process to make the correct mappings.

Another difference between point-to-point and multipoint interfaces is that while a point-to-point subinterface will change its protocol status to down when the assigned DLCI is not active, while for a multipoint subinterface or a physical interface, the protocol status is not affected by the PVC status.

## Disabling Inverse ARP

Can I disable Inverse ARP? Not entirely. You can disable the sending of ARP requests but you cannot disable responding to ARP requests. This is even more trickier than it sounds. If you disable ARP response on your end but the other end of the virtual circuit sends an ARP request to you, your router will respond to it. The other router receives the information and adds it to it’s dynamic mappings, but you could be surprised to find out that even your own router used the information it received and added the other end’s address to its dynamic mappings. So you disabled ARP but ARP still works. It must be disabled on both ends to stop working completly. Inverse ARP can be disabled on an interface, on a PROTOCOL or on a DLCI using the following command:

```
R(config-if)#no frame-relay inverse-arp {PROTOCOL [DLCI]}
```

Another option for disabling Inverse ARP on just some DLCIs is more of a trick. You assign those DLCIs to a subinterface that does not run IP. When does an interface not run IP? When there is no IP address assigned to it.

## Final thoughts

There is a lot more to talk about Frame Relay, and I will post more tips and tricks, but I think this is enough for a brief introduction. Here are a few things to remember:

* Frame Relay is a Layer 2 protocol. This means that the Frame Relay Cloud is transparent to the **traceroute** command
* Look for the status of the PVC – this will indicate if there is a configuration issue – using **show frame-relay pvc**
* Issue the **show frame-relay map** command to see the Layer 2 to Layer 3 mappings. Without proper mapping encapsulation will fail
* Frame Relay is a non-broadcast network so you have to pay special attention when dealing with broadcast and multicast traffic. Try to **ping 255.255.255.255 repeat 1** to see that everybody is responding.


# Frame Relay Encapsulations – IETF vs Cisco

When configuring a Frame Relay interface, the first thing you should do is enable Frame Relay Encapsulation. This is usually a fairly simple task, just use the classic **encapsulation frame-relay** command and it will work. Actually there are 3 options:

```
R(config)# interface serial1/0
R(config-if)#encapsulation frame-relay ?
 MFR Multilink Frame Relay bundle interface
 ietf Use RFC1490/RFC2427 encapsulation
 <cr >
 
```

The first one is [MultiLink Frame Relay](https://nyquist.eu/multilink-frame-relay/) but this deals with a particular architecture, so the primary options are **ietf** and **cisco**. Cisco is selected by default when you just hit enter.

## IETF vs Cisco Encapsulation

Encapsulation is the process of taking the upper layer information and wrapping it up with a new header and trailer before sending it to a lower layer. When we talk about Frame Relay, which is a Layer 2 protocol, the lower layer is the Physical Layer, so it is practically the first/last logical layer in the OSI stack

![Frame Relay Header](/files/1WLZmUFhqR62wMINfrfr)

On the ~~left~~  top you can see how the IETF encapsulation looks like, according to [RFC1490](https://tools.ietf.org/html/rfc1490)/[RFC2427](https://tools.ietf.org/html/rfc2427). The Frame is delimited at the beginning and at the end with an 8 bits **Flag** that is always set to 0x7E.

In the next 16 bits (can grow up to 32 bits) is the **Q.922 Address** portion of the Frame. You will see the value of this field in several show commands on Cisco routers. Inside the Q.922 Address, there is also the DLCI.

Next is the **Q.922 Control** field. This field is missing in the Cisco encapsulation.

Next there is an optional **Padding** field. When used, it has an 8 bit value of all zeros with the goal of aligning the data field to start exactly after a multiple of 16 bits chunks.

The **NLPID** field is used to identify the upper layer protocol . Cisco encapsulation also has a **Type** field that has the same function but the values used to identify the protocols are different. For example, IPv4 is identified as 0xCC by IETF and 0x0800 by Cisco.

Next is the **Data** field which has a variable size.

At the end we have a **FCS** field used for error checking and again the **Flag** Field that signals the end of the frame.

### 2. The Q.922 field – Here’s DLCI!

The Q.922 field, also known as the Q.922 Address contains the DLCI along with other options for the particular frame

![Frame Relay Q922](/files/lXmMZzcVC2yiuhgvInNp)

The **DLCI** is split between the 2 Bytes of the Q.922 field, in two chunks of 6 and 4 bits. That makes up for 10 bits, so the DLCI values that could be used theoretically should be in the range 0 – 1023. In practice, depending on the type of LMI used, only a subset of these can be used, while others are reserved.

The **C/R** field in the first Byte is used to signal if the frame is a Command or a Response.

The **EA** field is also split and it signals the possibility of expanding the address field with up to 2 additional bytes, offering a bigger address space.

The **FECN** and **BECN** fields are used for the mechanism of Explicit Congestion Notification. Frames marked with the FECN bit signal a congestion in the downstream direction, that is from the sender to the receiver, while frames marked with the BECN bit signal a congestion in the upstream direction, that is from receiver to sender.

The **DE** bit stands for Discard Eligible and the frames marked with it can be discarded if there is a congestion along the way. Usually, when buying a Frame Relay service, the provider offers a CIR value – Committed Information Rate, that is guaranteed to be delivered. If the customer sends data at a rate that is over the CIR value, the offending frames are not discarded, but they are marked with the DE bit and could be discarded in the event of a congestion on the path to the destination.

### 3. Final thoughts

In theory, the encapsulation should be matched between DTE devices at the end of the PVC. Also in theory, **ietf** encapsulation should be used when connecting Cisco devices to non-Cisco devices. In real life though, there are rarely any problems with this because many other devices support **cisco** encapsulation, but also because Cisco devices work even with a mismatched encapsulation. They will send packets with the defined encapsulation but will accept packets with either encapsulation.


# Multilink Frame Relay

Frame Relay can be used to group multiple physical interfaces in one logical interface, much like an EtherChannel interface available on the Ethernet Switches. It offers load-balancing and higher availability since the link will be up as long as one physical connection is up.

In order to do this, bothe the DTE and DCE device must be configured to support this configuration. As you will see, this is not an end-to-end configuration, but rather a local config, between the DTE and the local DCE. The DCE will switch the frames to the other DTEs just as with normal Frame Relay encapsulation

![MultiLink Frame Relay](/files/BiRcpS66TM2nzlW0OjBh)

## First DTE

```
!On R1:
R1(config)# interface mfr1
R1(config-if)# ip address 123.0.0.1 255.255.255.0
R1(config-if)# exit
R1(config)# interface serial1/0
R1(config-if)# encapsulation frame-relay mfr 1
R1(config-if)# no shut
R1(config-if)# exit
R1(config)# interface serial1/1
R1(config-if)# encapsulation frame-relay mfr 1
R1(config-if)# no shut
```

## Frame Relay Switch

```
!On FRS:
FRS(config)# frame-relay switching
FRS(config)# interface mfr1
FRS(config-if)# frame-relay intf-type dce
FRS(config-if)# exit
FRS(config)# interface serial1/0
FRS(config-if)# encapsulation frame-relay mfr 1
FRS(config-if)# no shut
FRS(config-if)# exit
FRS(config)# interface serial1/1
FRS(config-if)# encapsulation frame-relay mfr 1
FRS(config-if)# no shut
FRS(config-if)# exit
! FRS connects to R2 overs serial 1/2
FRS(config)# interface serial1/2
FRS(config-if)# encapsulation frame-relay
FRS(config-if)# frame-relay intf-type dce
FRS(config-if)# no shut
FRS(config-if)# exit
! Now let's glue them together:
FRS(config)# connect R1_R2 mfr1 102 Serial1/2 201
```

## Second DTE

```
!On R2
R2(config)# interface serial1/2
R2(config-if)# ip address 123.0.0.2 255.255.255.0
R2(config-if)# encapsulation frame-relay
R2(config-if)# no shut
```

## Verification

The logical MFR1 interface can be used just like a physical frame-relay interface. It can also be set up with subinterfaces. By default, dynamic Inverse ARP mappings will be assigned to the Multilink interface:

```
R1#sh frame-relay map
MFR1 (up): ip 123.0.0.2 dlci 102(0x66,0x1860), dynamic,
              broadcast,, status defined, active
```

To verify the status of a Multilink bundle use:

```
R1#sh frame-relay multilink 
Bundle: MFR1, State = up, class = A, fragmentation disabled
 BID = MFR1
 Bundle links:
  Serial1/0, HW state = up, link state = Up, LID = Serial1/0
  Serial1/1, HW state = up, link state = Up, LID = Serial1/1
```


# Frame Relay Switching

## Frame Relay Switching

When setting up a Frame Relay DCE device, we need to follow the following steps.\
1\. Enable Frame Relay Switching\
2\. Change The interface type\
3\. Glue the DLCIs together

### Enable Frame Relay Switching

The command is ultra simple, but without it the device will try to decapsulate the frames and look up the Layer 3 information. This is not needed on a Frame Relay Switch.

```
FRS(config)# frame-relay switching
```

### Change the interface type

Interfaces connected to DTE routers must be set to DCE to enable LMI

```
FRS(config-if)# frame-relay intf-type dce
```

On interfaces connecting to other Frame Relay switches, the type must be set to nni – Network-to-Network Interface

```
FRS(config-if)# frame-relay intf-type nni
```

### Glue the DLCIs together

There are 2 methods for connecting a DLCI on an interface to a DLCI on another interface. The old method needed two commands, one for each interface, like this:

```
FRS(config)# interface INTERFACE-1
FRS(config-if)# frame-relay route DLCI-1 interface INTERFACE-2 DLCI-2
FRS(config-if)# exit
FRS(config)# interface INTERFACE-2
FRS(config-if)# frame-relay route DLCI-2 interface INTERFACE-1 DLCI-1
```

To verify, use:

```
FRS#sh frame-relay route
Input Intf 	Input Dlci 	Output Intf 	Output Dlci 	Status
Serial1/1       102 		Serial1/2       201 		inactive
Serial1/2       201 		Serial1/1       102 		inactive
```

The newer method needs just one command in the global config:

```
FRS(config)# connect CONN-NAME INTERFACE-1 DLCI-1 INTERFACE-2 DLCI-2
```

and to verify, use:

```
FRS1#sh connection all

ID   Name               Segment 1            Segment 2           State       
========================================================================
1    R1R2              Se1/1 102            Se1/2 201            UP          
```

## Switching over a Tunnel interface

There are situations where you have Frame Relay routers as Provider edge, but use another technology as the Provider Core. You can still perform Frame Relay Switching by tunneling the Frames from one PE router to another

![](/files/BxQ9RJ1g7m2Cgxtzjopr)

\
You must configure the frame relay interfaces similarly on the two Frame Relay switches and then configure the Tunnel interface on each router:

```
FRS(config)# interface TUNNEL0
FRS(config-if)# tunnel source {INTERFACE|SRC-IP-ADDR}
FRS(config-if)# tunnel destination {INTERFACE|SRC-IP-ADDR}
FRS(config-if)# ip {unnumbered INTERFACE| address TUN-IP-ADDR}
```

Then create the frame-relay routes, making sure you use the same DLCI on the Tunnel interface on both routers:

```
!On FRS1
FRS1(config)# interface INTERFACE-1
FRS1(config-if)# frame-relay route DLCI-1 interface TUNNEL0 DLCI-TUN
!On FRS2
FRS2(config)# interface INTERFACE-2
FRS2(config-if)# frame-relay route DLCI-2 interface TUNNEL0 DLCI-TUN
```

You cannot use the connect command to create the routes.


# Routing over Frame Relay

## Topologies

### Full Mesh

![Frame Relay Full Mesh Topology](/files/UimUd4gri4A4jLQh0YW8)

The simplest Frame Relay topology is the Full Mesh topology, where each router has a dedicated virtual circuit to another router. Unfortunately this design is rarely found in real life because each additional circuit costs. Of course, having so many circuits available makes it easy to use any kind of links. The simplest method here is to use multipoint interfaces with different mappings (static or dynamic) for each router, but with all of them in the same subnet. Of course, you can also use point-to-point sub-interfaces with different subnets, but this is not a scalable solution.

### Hub and Spoke

In a hub and spoke topology we have one central router with connections to all other routers. The spokes do not have direct circuits between each other, they can only communicate via the Hub. To make things work you can use several point-to-point interfaces on the Hub with a different subnet on each link to a spoke, or you can use a multipoint interface with a single subnet for all routers

![Frame Relay Hub and Spoke](/files/DF415qnsbdVZkXyTM9z5)

While it is more cost effective, this topology introduces several challenges, like:

* **Mapping** – On each spoke, the inverse ARP will only resolve the Hub. Therefore static mapping for the other hosts is needed in order to enable communication between the spokes.
* **Broadcasts** – The hub will have to route between the 2 DLCIs in order to pass traffic from one spoke to another. This means that broadcast traffic will not be forwarded between spokes.
* **TTL** – When routing on the bridge TTL is decremented. This will be a problem when sending traffic with TTL=1 as in OSPF or eBGP. One option is to create a bridge between interfaces in order not to route the packet and to avoid TTL decrementation.

## IPv4 Routing Issues

The following considerations regard a Hub-and-Spoke topology with the hub using a multipoint subinterface. This topology has the most issues when trying to run a routing protocol over it. Using several point-to-point subinterfaces simplifies the connectivity model. Also, a multipoint subinterface works just like a physical Frame Relay subinterface, but it needs additional configurations because LMI assigns the DLCIs to the physical interface and so it helps automatic detection for Frame Relay mappings using Inverse ARP. Physical interfaces also have split-horizons for RIP disabled by default, which helps convergence. On Subinterfaces it has to be manually disabled. Split Horizons for EIGRP is enabled by default on all interfaces.

### RIP

1. RIPv1 uses broadcasts and RIPv2 uses multicast to send routing updates, so first of all, the DLCIs have to be able to transport broadcasts.
2. Since updates are sent as broadcast or multicast, the updates sent by one spoke will reach the hub but will not reach the other spokes. (Broadcast/Multicast traffic si not routed). Solutions:
   * The hub should forward the updates to all spokes, but the Split Horizons rule prevents this. It should be disabled using **no ip split-horizon** at the (sub)interface level. The next-hop address for all routes will point to the Hub’s IP address
   * If Split Horizon cannot be disabled, then we should send the updates as unicast from one spoke to another. This requires that each spoke declares its neighbors with **neighbor NEIGH-IP** command inside the routing process. The next-hop for spoke routes will point to each spoke’s IP address so we will also need mappings for the neighbors address on the (sub)interface.
   * You could also create a tunnel interface between the spokes but this will turn the topology into a Pseudo-Full-Mesh
3. Cisco routers send RIP messages with TTL=2, so they will not be affected by a TTL decrement on the Hub.

### EIGRP

1. EIGRP uses multicast to send routing updates, so first of all, the DLCIs have to be able to transport broadcasts.
2. Since updates are sent as broadcast or multicast, the updates sent by one spoke will reach the hub but will not reach the other spokes. (Broadcast/Multicast traffic si not routed). Solutions:
   * The hub should forward the updates to all spokes, but the Split Horizons rule prevents this. It should be disabled using **no ip split-horizon eigrp** at the (sub)interface level. The next-hop address for all routes will point to the Hub’s IP address
   * By default, EIGRP will change the advertised next-hop address to the outgoing interface. This is why the spokes will have all EIGRP learned routes pointing to the Hub’s IP Address. This behavior can be changed if we configure the hub with:

     ```
     R(config-if)# no ip next-hop-self eigrp AS-NUMBER
     ```

     Now, the routes will point to each originating spoke’s IP address. We will then need frame-relay maps pointing to the spokes.
   * If Split Horizon cannot be disabled, then we should send the updates as unicast from one spoke to another. This requires that each spoke declares its neighbors with **neighbor NEIGH-IP INTERFACE** command inside the routing process. The next-hop for spoke routes will point to each spoke’s IP address so we will also need mappings for the neighbors address on the (sub)interface. Unlike RIP, when a neighbor is defined in EIGRP, it will disable multicast updates on that interface. We will have to define all spokes and the hub as neighbors with each other
   * You could also create a tunnel interface between the spokes but this will turn the topology into a Pseudo-Full-Mesh
3. Cisco routers send EIGRP messages with TTL=2, so they will not be affected by a TTL decrement on the Hub.

### OSPF

OSPF uses the concept of network type, which defines its mode of operation. By default, on Frame Relay physical interfaces and multipoint subinterfaces, the network type used is NON\_BROADCAST, while on point-to-point subinterfaces it is POINT\_TO\_POINT.

#### **Non Broadcast**

1. **Define neighbors**: Since it is considered a non-broadcast medium, all packets are expected to be sent as unicast so the neighbors must be statically defined. We will need mappings for each spoke on the hub, but the DLCIs don’t have to support broadcasts.
2. **Force the Hub as the DR:**&#x4F;n NON\_BROADCAST networks a DR is elected. Since adjacencies are only formed with the DR, it should be forced on the Hub. To force the DR on the hub, use:

   ```
   R(config-if)# ip ospf priority 0
   ```

   on each spoke (sub)interface.
3. **Hellos/Dead timer: 30/120 sec**
4. Because OSPF packets are sent with TTL=1, spokes cannot become neighbors
5. The next hop address for the routes advertised by spokes, are set to the spoke’s IP address, therefore static mappings are required on each spoke, pointing to the other spokes.

#### **Point to Multipoint**

A better option for Frame Relay is to use the POINT\_TO\_MULTIPOINT network type. This type of network has the following characteristics:

1. **Neighbor auto-discovery:** Packets are sent as multicast so neighbors can be auto-discovered, but the DLCIs must support broadcasts
2. **No DR is elected**
3. **Hellos/Dead timer: 30/120 sec**
4. Because OSPF packets are sent with TTL=1, spokes cannot become neighbors
5. The next hop address for the routes advertised by spokes is changed to the hub’s address, so routes on the spokes will point to the hub. No need for additional mappings

#### **Point to Multipoint Non Broadcast**

The third option is the POINT\_TO\_MULTIPOINT\_NON\_BROADCAST network type. This network type is a mix of the POINT\_TO\_MULTIPOINT and the NON\_BROADCAST network types.

1. **Define neighbors**: Packets are sent as unicast so static nieghbors must be defined. Since traffic is unicast, DLCIs do not have to support broadcasts.
2. **No DR is elected**
3. **Hellos/Dead timer: 30/120 sec**
4. Because OSPF packets are sent with TTL=1, spokes cannot become neighbors
5. The next hop address for the routes advertised by spokes is changed to the hub’s address, so routes on the spokes will point to the hub. No need for additional mappings

#### **Point to Point**

This option is used by default on point-to-point subinterfaces. You can’t use this network type in a hup and spoke architecture because the hub will be expecting only one neighbor on the interface. However, you can have the hub configured to be POINT\_TO\_MULTIPOINT and the spokes to be POINT\_TO\_POINT (sub)interfaces as long as you configure the same HELLO and DEAD timers on each interface.

1. **Neighbor auto-discovery:** Packets are sent as multicast so the neighbor can be auto-discovered, but the DLCIs must support broadcasts. On point-to-point networks, you can’t define static neighbors.
2. **No DR is elected.**
3. **Hellos/Dead timer: 10/40 sec**
4. The next hop address for the routes advertised by spokes is changed to the hub’s address, so routes on the spokes will point to the hub. No need for additional mappings

#### **Broadcast**

Normally, you would not use this type of network on Frame Relay interfaces, but it will still work.

1. **Neighbor auto-discovery:** Packets are sent as multicast so the neighbor can be auto-discovered, but the DLCIs must support broadcasts. On point-to-point networks, you can’t define static neighbors.
2. **DR is elected**: Force it on the hub
3. **Hellos/Dead timer: 10/40 sec**
4. The next hop address for the routes advertised by spokes, are set to the spoke’s IP address, therefore static mappings are required on each spoke, pointing to the other spokes.

#### BGP

* BGP uses unicasts so there is no need for the DLCI to allow broadcasts.
* eBGP sends data with TTL=1. The solution is to use eBGP multihop.
* iBGP doesn’t change the next-hop when forwarding routes, so we might need to statically map the other spokes to the DLCIs.

## IPv6 Routing Issues

### RIPng

1. Fist, updates are sent as multicast, so DLCIs must allow multicasts to be sent.
2. When seding updates, RIPng uses the Link-local address as the source, so all routes point to the link-local addresses. You must statically define mappings for these addresses in order to be able to resolve the next-hop layer 2 address. You will probably have to disable the RIP split Horizon rule, using:

   ```
   R(config)# ipv6 router rip PROCESS-NAME
   R(config-rtr)# no split-horizons
   ```
3. You can’t define static neighbors with RIPng
4. If you use multiple subinterfaces, you might see that the same link-local address is used for all subinterfaces of a serial link. This makes it impossible to map each destination to the DLCI it shoud use (since the same destination si already mapped on another DLCI). The solution is to manually define the link-local address.

### EIGRP for IPv6

1. Fist, the EIGRP process must be enabled with:

   ```
   R(config)# ipv6 router eigrp AS-NUMBER
   R(config-rtr)# no shutdown
   ```
2. Then, updates are sent as multicast, so DLCIs must allow multicasts to be sent.
3. When seding updates, EIGRP for IPv6 uses the Link-local address as the source, so all routes point to the link-local addresses. You must statically define mappings for these addresses in order to be able to resolve the next-hop layer 2 address. You will probably have to disable the EIGRP split Horizon rule, per interface, using:

   ```
   R(config-subif)# no ipv6 split-horizon eigrp AS-NUMBER
   ```
4. You can define static neighbors with EIGRP for IPv6, using:

   ```
   R(config-rtr)# neighbor NEIGH-IPV6 INTERFACE
   ```

   but it will also disable multicast updates, so you will have to define neighbors on all routers
5. If you use multiple subinterfaces, you might see that the same link-local address is used for all subinterfaces of a serial link. This makes it impossible to map each destination to the DLCI it shoud use (since the same destination si already mapped on another DLCI). The solution is to manually define the link-local address.

### OSPFv3

The same considerations from IPv4 apply to IPv6. The default network type is NON\_BROADCAST. The network type can be changed with:

```
R(config-if)# ipv6 ospf network {broadcast|non-broadcast|point-to-point|point-to-multipoint [non-broadcast]}
```

1. When seding updates, OSPF for IPv6 uses the Link-local address as the source, so all routes point to the link-local addresses. You must statically define mappings for these addresses in order to be able to resolve the next-hop layer 2 address.
2. You can define static neighbors with OSPFv3, per interface, using:

   ```
   R(config-subif)# ipv6 ospf neighbor NEIGH-IPV6 ...
   ```
3. If you use multiple subinterfaces, you might see that the same link-local address is used for all subinterfaces of a serial link. This makes it impossible to map each destination to the DLCI it shoud use (since the same destination si already mapped on another DLCI). The solution is to manually define the link-local address.

### MP-BGP

MP BGP supports IPv6 in a similar way as with IPv4. You just have to enable the [IPv6 Address Family](https://nyquist.eu/ipv6-routing/#61_Start_MP-BGP_for_IPv6). As a neighbor address you can use a unicast or a link-local address, but you will have to check if the correct mappings exist (no need for broadcast support, though).


# Bridging


# Bridging on a router

## Bridging

Transparent Bridging is the default operational mode of switches. They bridge between interfaces and switch between them without modifying any data in the frames.\
Routing is the default operation mode of routers. They route between interfaces, and when doing this they modify the packets (Source and Destination MAC, TTL, etc)\
To enable routing, you use:

```
Sw(config)# ip routing
```

This command is default on routers, but must be manually entered on switches.\
Now, to enable bridging on routers, you have a few options:

### Transparent Bridging

By default, a router can only route a protocol. It can also bridge it, but it cannot do both at the same time. In order to enable transparent bridging, you will have to disable routing:

```
R(config)# no ip routing
```

By disabling routing, we will not be able to route between interfaces, but the router can still act as an IP host. Besides bridging interfaces, you can also set an IP address on the physical interfaces to make the router accessible in that subnet. One workaround to having routing disabled is to set the same IP address on all interfaces that are part of the bridge, thus making the router accessible on all interfaces.

### CRB – Concurrent Bridging and Routing

CRB is a way of performing both bridging and routing at the same time on a router. In this way, some interfaces can be used for routing, while others will be bridged, but they cannot do both at the same time.\
To enable CRB, leave ip routing on and set:

```
R(config)# bridge crb
```

Next you will have to define [how interfaces are bridged](broken://pages/Ws1P2OI8dJK0VWR8uqux).

### IRB – Integrated Bridging and Routing

IRB is an upgrade from CRB, where you can do both routing and bridging at the same time. When using IRB, you can define a Bridged Virtual Interface (BVI) that will be able to route the traffic on the bridged interfaces, just like an SVI on a switch.\
To enable irb, leave ip routing on and set:

```
R(config)# bridge irb
```

You will still have to define [how interfaces are bridged](broken://pages/Ws1P2OI8dJK0VWR8uqux).\
Then, in order to set up the BVI interface, first define what protocols can be routed:

```
R(config)# bridge BRIDGE route {ip|clns}
```

Then, you can use the BVI interface, just like a normal routed interface:

```
R(config)# interface bvi BRIDGE
R(config-if)# ip address IP-ADDR MASK
```

## Bridging interfaces

After defining what type of bridging to use, you have to define how the interfaces are bridged. Follow these steps:

### Define the Spanning Tree Protocol

First we need to define what type of Spanning Tree will run on this bridge:

```
R(config)# bridge BRIDGE protocol {ieee|dec|ibm|vlan-bridge}
! ieee = 802.1D
```

### Assign interfaces to the bridge

```
R(config-if)# bridge-group BRIDGE
```

At this point, you should check that the interfaces in this group act as a Layer 2 Switch:

```
R5#sh bridge group
Bridge Group 1 is running the IEEE compatible Spanning Tree protocol
   Port 4 (FastEthernet0/0) of bridge group 1 is forwarding
   Port 5 (FastEthernet0/1) of bridge group 1 is forwarding
```

To see a list of MAC addresses learned on each bridge, use:

```
R# show bridge [BRIDGE]
```

and to see the status of the spanning tree, use:

```
R# show spanning-tree [brief]
```

## Bridging with Frame Relay interfaces

### Different Encapsulations

When connecting an Ethernet interface (R1-Fa0/1) to another Ethernet interface(R2-Fa0/1) that is part of a bridge(R2-BVI1), you can ping from one side (R1-Fa0/1) to the other (R2-BVI1). When connecting a Frame Relay interface(R1-S1/0) to another Frame Relay interface(R2-S1/0) that is part of a bridge(R2-BVI1), you will encounter a problem with encapsulation. On one side (R1-S1/0) sends packets with Frame Relay encapsulation, and on the other side (R2-BVI1), packets with an Ethernet-ARPA encapsulation are expected. The only solution here is to create one bridged interface on each side (R1-BVI1, R2-BVI1).

### Mapping

Another problem arises when using multipoint frame-relay interfaces. On a point-to-point interface that is part of the bridge, the router will use the DLCI assigned to the interface to send data. But on multipoint interfaces, an explicit mapping is required for each DLCI:

```
R(config-if)# frame-relay map bridge DLCI [broadcast]
!broadcast needed to send BPDUs
```

### Switching between spokes

It won’t work. If the interface on the hub is a multipoint interface, then it will consider the spokes as two hosts connected on the same bridge port. In this case, it will drop the frame, considering that the frame sent by the source also reached the destination. There are 2 options here: either use a tunnel or force the traffic to go on another link by manipulating the spanning-tree (provided you have another link).


# MTU 101

MTU stands for Maximum Transmission Unit. This is the amount of data that can be transmitted by one protocol. MTU is used at every layer of the OSI stack, but it’s value is closely related to the layer/protocol.

## On a router

### Layer 2 – mtu

On a router, the mtu command applied on an interface defines the value of the L2 payload. This means how much data a L2 frame can contain, without the L2 header or trailer.\
By default, this value is 1500 Bytes on Serial and FastEthernet interfaces, but it can be changed with:

```
R(config-if)# mtu BYTES
! Default: 1500
```

To see the current value used as the L2 MTU, use:

```
R# show interface INTERFACE | i MTU
  MTU 1500 bytes, BW 10000 Kbit, DLY 1000 usec,
```

As you will see in the next section, on a Cisco device the **mtu** and the **ip mtu** values refer to the same portion of a packet. The value defined for the L2 MTU is closely related to the value defined for the L3 MTU because the L3 MTU can’t be larger than the L2 MTU. When setting the L2 MTU to a value lower than L3 MTU, the L3 MTU also changes. When setting the L3 MTU you can’t set a value larger than L2 MTU, but it can be lower. The L2 MTU doesn’t change.\
When a router receives a frame that is larger than the L2 MTU (comparing only the L2 payload), the packet is dropped.

### Layer 3 – ip mtu

On a router, the ip mtu command applied on an interface defines the size of the L3 packet, including its headers. As you can see, this represents the same portion as the one used as the L2 MTU. By default, the L3 MTU value for Serial or FastEthernet connections is 1500 Bytes, same as the default L2 MTU. You can change it with:

```
R(config-if)# ip mtu BYTES
! Default: 1500
```

To see the current value used as the L3 MTU, use:

```
R# show ip interface INTERFACE | i MTU
  MTU is 1500 bytes
```

If you need to send L3 packets that are larger than the IP MTU, fragmentation occurs. The router will split the original packets in packets that can be accommodated by the L3 MTU size. Well, this works as long as the packet doesn’t have the DF(Don’t Fragment) bit set. If such a packet arrives and the L3 MTU is smaller than the packet size, then the packet is dropped and an ICMP message is generated (packet too big)

### Layer 4 – ip tcp mss

On a router, the MSS (Maximum Segment Size) represents the MTU value for TCP – How much data can a TCP packet contain. This value is negotiated during TCP handshake and is chosen as the smallest value configured on the sender or the receiver.\
When a Cisco Router is a sender or a receiver, it uses a MSS value of 1460 Bytes for local destinations (same subnet) and 536 bytes for remote destinations. This value can be changed with:

```
R(config)# ip tcp mss BYTES
```

When a router is neither the sender, nor the receiver it can intercept TCP packets and modify the MSS value to prevent dropping larger packets than the network would accept. To enable this, use:

```
R(config)# ip tcp adjust-mss BYTES
```

## On a switch

### Layer 2 – system mtu

On a switch you can’t configure individual MTU values for every port. Instead, a global MTU value is configured that is used on all ports. This value only changes when the router reloads.\
The definition for the L2 MTU is the same as for routers and has the same default value: 1500 Bytes. It can be changed with:

```
Sw(config)# system mtu BYTES
! Default 1500. Max:1998. Requires reload
```

When a switch receives frames that are larger than the L2 MTU, it drops them.

```
Sw# show system mtu
System MTU size is 1500 bytes
System Jumbo MTU size is 1550 bytes
Routing MTU size is 1500 bytes.
```

#### **Layer 2 – system mtu jumbo**

For Gigabit interfaces, the value for Jumbo MTU is used. To configure the switch to support Jumbo frames on Gigabit ports, use:

```
Sw(config)# system mtu jumbo BYTES
! Default 1500. Max: 9000. Requires reload
```

If a switch receives frames larger than the JUMBO MTU, it drops them.\
If a switch receives a frame larger than the L2 MTU, but smaller than the JUMBO MTU, it will still drop them if they are destined for a FastEthernet port, regardless of the port they were received on.

```
Sw# show system mtu
System MTU size is 1500 bytes
System Jumbo MTU size is 1550 bytes
Routing MTU size is 1500 bytes.
```

### Layer 3 – system mtu routing

This value represents the L3 MTU size and can’t be bigger than the L2 MTU size. If the L2 MTU size is set lower than the L3 MTU size, the L3 MTU size will also change after reload.\
This value is also used by OSPF when advertising its MTU. To change it, use:

```
R(config)# system mtu routing BYTES
! Default 1500. Doesn't require reload
```

To verify, use:

```
Sw# show system mtu
System MTU size is 1500 bytes
System Jumbo MTU size is 1550 bytes
Routing MTU size is 1500 bytes.
```

## When to worry about MTU

Usually, the default MTU size is not modified, unless the devices work in an environment where additional encapsulation is used. Additional encapsulation occurs in the following situations:

* L2 frames encapsulated in other L2 frames
  * **PPPoE**: the L3 MTU of the PPPoE interface should be set to 1492.
  * **Q-in-Q Tunneling**: The L2 MTU should be set to 1504
* L3 packets are encapsulated in other L3 packets
  * **GRE Tunneling**: Adds 24 Bytes to the IP Header, so when crafting packets, the router will limit the L3 size to MTU-24 Bytes, in order to accomodate the packets on the network. Fragmentation may occur
  * **Other L3 Tunnels**
* **MPLS** – adds 4 bytes for each label. For MPLS L3 VPN, there are 2 labels, so 8 Bytes. E.g. if L2 MTU is 1500, L3 packets larger than 1492 Bytes will be fragmented (or dropped). Some hardware allow you to set an **MPLS MTU** that is larger than L2 MTU, in order to reduce fragmentation. (See Baby Giants)

The problems are not always obvious because only packets that are very close in size to the maximum MTU, when adding extra encapsulation will become larger in size than the allowed MTU. Usually, normal pings are smaller and don’t get dropped. A better way to test it is to use large ping packets with the DF bit set:

```
R# ping IP-ADDR size 1500 df-bit
```

Actually, this is the method used by the **Path MTU Discovery** (PMTUD) process. In order to avoid fragmentation, which adds additional delay, a host that supports PMTUD can discover what is the largest MTU on the path to the destination, and only send packets large enough not to be fragmented or dropped. You can enable this functionality when a router is the sender with:

```
R(config)# ip tcp path-mtu-discovery [age-timer TIMER|infinite]
```


# Wireless


# Wireless Principles

## RF Spectrum

A radio wave is an electromagnetic field radiating from a sender. The wave propagates towards a receiver which receives its energy. Electromagnetic waves are caracterized by their wave length. The length of the wave is defined as the physical distance the wave covers in one cycle so it is $$\lambda ={\frac {v}{f}},,,$$wher v is the speed of the wave and f is the frequency of the wave. Since Electromagnetic waves travel at the speed of light $$c = 3 x10^8 m/s$$, the wavelength is $$\lambda ={\frac {c}{f}},$$  . For example an electromagnetic (radio) wave of 100 MHz has a wave length of&#x20;

$$
\lambda ={\frac {3\times 10^8 m/s}{100 \times 10^6 Hz} } =  {\frac {3 \times 10^8 m/s}{1 \times 10^8 1/s}} = 3m
$$

since $$1 Hz = \frac{1}{s}$$ and represents the frequency of the wave.

A wave also has an amplitude which represents the ammount of energy that is injected in one cycle. The amplitude of a wave can be increased through amplification. Amplification can be active (more energy is applied) or passive (energy is focused with an antenna). Decreasing the amplitude is called attenuation. While travelling further from the source wave amplitude suffers attenuation.As energy is absorbed by other obstacles on the path, attenuation can happen much faster depending on the environment. But there are also reasons for weaker signals at receivers which make up the "Free Path Loss"

* The signal is snet in all directions so the sender's energy is spread in all directions. With antennas energy can be focused in certain areas but there is no perfect way of focusing the energy from sender to receiver
* The receiver has a certain size and can only collect a limited ammount of the energy that is sent.

Regulations are in place to determine the (maximum) ammout of power that should be used for each device type based on the distance where the signal is expected to be sent.

## RSSI and SNR

To determine how much of the original signal reaches the reciever we can use RSSI (Received Signal Strenght Indicator) and SNR (Signat to Noise Ratio). RSSI - Noise = SNR

RSSI calcualtion is not easy because the reciever doesn't know how much power was originally sent so RSSI is a relative value obtained by comparing received packets to each other. Each vendor has it's own ranges so the same RSSI value from one vendor can't be compared with the RSSI value from another vendor. For Cisco, a good RSSI value is higher than -67dBm.

RCPI (Received Channle Power Indicator) is an attempt to standardize RSSI across vendors.

The noise represents the interferences in your noise level so lower is better (e.g -95dBm).&#x20;

SNR represents the ratio between Signal and Noise but can be obtained with the formula SNR = RSSI - Noise. Since RSSI and Noise are negative values in dBm the result should end up positive in most situations. A SNR>20dBm is good.

SINR (Signla to Interference plus Noise Ratio) = RSSI - (interference + noise). A SINR > 25dBm is required for voice over wireless

## Decibels and Watts

Power is measured in Watts but power calculations can be a complex excercise. For this reason, a simpler apporach is to use the logarithmic measurement that expresses the ammount of power relative to a reference, which is called decibel (dB)

When the reference power is equal to the compared power, the dB difference between them is 0. In addition, these shortcuts can be used to quickly evaluate the power value based on it's reference

* +10 dB means the compared value is 10 times more powerful than the reference.
* \+ 3 dB means the compared value is 2 times more powerful than the reference
* -3 dB means the compared value is half the power of the reference
* -10 dB means the compared value is 1/10 of the power of the reference.

Sometimes the decibel value also has a reference indicator:

* dBm - the reference value is a power of 1mW
* dBd - the reference value is a [dipole antenna](https://en.wikipedia.org/wiki/Dipole_antenna)&#x20;
* dBi - the reference value is an isotropic antenna (an omnidirectional antenna)

EIRP (Effective Isotropic-Radiated Power) \[dBm] = TX \[dBm] + Antenna\[dBi] + Cable Loss\[dB]

## Antenna Characteristics

Antennas fall in 2 main categories:

* Omnidirectional: They radiate a signal with the same strength in all directions so the energy is evenly distributed
  * Dipole
* Directional: the directional antenna radiates most of the energy in some direction so it is not evenly distributed. Because they are stronger in some areas they add "gain".
  * Yagi
  * Patch

The term directional here refers to all directions in a 3D space but vendors typically provide 2x 2D representations for the Azimuth (horizontal) and the elevation plane.


# Wireless Implementations

## Wireless Standards

{% tabs %}
{% tab title="802.11 a/b/g (Legacy)" %}
Year ratified:&#x20;

* 1999 (a/b)
* 2003 (g)

Frequency Band:&#x20;

* 5 GHz (a)
* 2.4 GHz (b,g)

Data Rates:&#x20;

* 11Mbps (b)
* 54Mbps (a,g)

Features:

* SISO
  {% endtab %}

{% tab title="802.11n" %}
Year ratified:&#x20;

* 2009

Frequency Band:&#x20;

* 5 GHz
* 2.4 GHz

Data Rates:&#x20;

* Up to 600 Mbps (channel bonding for up to 40MHz)

Features:

* backwards compatible with 802.11a/b/g
* MIMO
  {% endtab %}

{% tab title="802.11ac" %}
Year ratified:&#x20;

* 2013

Frequency Band:&#x20;

* 5 GHz&#x20;

Data Rates:&#x20;

* 1300 Mbps - Wave 1 (channel bonding of up to 80MHz)
* 6930 Mbps - Wave 2 (channel bonding of up to 160MHz)

Features:

* 802.11ac is backwords compatible with 802.11a and 802.11n
* MU-MIMO
  {% endtab %}

{% tab title="802.11ax (Wi-Fi 6)" %}
Year ratified:&#x20;

* 2021

Frequency Band:&#x20;

* 5 GHz
* 2.4 GHz&#x20;

Data Rates:&#x20;

* 4800 - Wave 1
  {% endtab %}
  {% endtabs %}

**SISO** is a system where a system uses a single antena at a time even if it has multiple antennas. Systems that can use multiple antennas symultenousley are called **MIMO**. MIMO incorporates 3 technologies:

* **MRC (Maximal Ratio Combining)** - a MIMO receiver uses MRC to combine energies from multiple recive chains
* **Beamforming** - a MIMO transmitter can coordinate the signal sent from each antenna so that the receiver gets a better signal. Cisco ClientLink is a beamforming technology
* **Spatial Multiplexing** - requires a MIMO transmitter and a MIMO receiver and allows the transmitter to split the data in multiple streams and send them to each antenna of the receiver.

While these features improve communication between one sender and one receiver at a time, **802.11ac MU-MIMO** allows the AP to transmit frames to multuple clients at the same time.

## Wireless Component Roles

### Clients and Access Points (APs)

An AP functions similarly to an Ethernet hun in that only one device can talk to the AP at a given time, over a shared media. A client's connection state to an AP can be one of:

* Not authenticated and not associated
* Authenticated but not associated (yet)
* Authenticated and associated - only in this state the data can flow

The associataion process has several steps:

1. mobile station sends a probe to discover available networks. Probe requests are sent to BSSID FF:FF:FF:FF:FF:FF (it will be received by all APs) and advertise the supported data rates and capabilities of the station
2. APs receving the probe request check to see if they support any of the advertised data rates and a probe response is sent with the SSID, supported data rates, encryption type and capabilities of the AP
3. Based on the responses received, the mobile station chooses a compatible network and sends an 802.11 authentication Open message with Seq set to 1(not the same authentication as WPA or 802.1x)
4. The AP receives the authentication frame and reponds with autehntication Open and Seq set to 2. (since authentication is open most requests should be succesful)
   1. If AP receives any frame other than an authentication or probe request from a station it will respond with Deauthentication frame and it will place the station in an "unauthenticated and unassociated state"
5. A station that received the Authentication Open with Seq=2 frame will send an association request to the AP.
   1. If AP receives any frame other than an assocation request from a station it will respond with Deassociation frame and it will place the station in an "authenticated but unassociated state"
6. If the asociation paramters match, the AP will create an Association ID and reply to the station with an Association response. At this point the client is authenticated and associated.

### Wireless controller

Enterprise solutions may require a large number of APs that would be difficult to adminsitrate and coordinate if they act as independent APs. For this reason there are solutions that make use of a Wireless Controller. Cisco's solution is called WLC (Wireless LAN Controller). In this case the functions of a traditional AP are split between the AP and the controller.

#### CAPWAP (Control and Provisioning of Wireless Access Points)

CAPWAP is an open protocol that enables a WLC to manage APs. The AP and WLC build a secure DTLS tunnel (control plane) to communicate. The client data is encapsulated with a CAPWAP header and is sent to the WLC.

#### Mobility Controller (MC) and Mobility Agent (MA)

MA and MC are functions running on WLC. MA is responsible to terminate CAPWAP tunnels so it maintains a cliend database while MC provides mobiloity management tasks including roaming, wireless IPS, guest access. MA reports local and roamed client states to MC.

#### POP and PoA FUnctions

The POP is the Point of Presence for the client. It anchors the client IP Address and is used for security policy applications. The PoA is the Point of Attachement. It moves with user AP connectivity and it is used for user mobility and QoS policy application.

Before a user roams the POP and PoA are the same but if the user roams the PoA may move as well.


# Wireless Roaming

## Client perspective

A wireless client decides to roam to a different AP when the connection to the current AP si degraded. The roaming decision is entirely on the client side and can be caused by:

* maximum retries exceeded: Each vendor has a different threshold. One threshold could trigger a shift to a lower data rate and another threshold coudl trigger roaming
* Low RSSI
* Low SRN
* Proprietary roaming parameters - In some scenarios the APs or the controllers can communicate with the clients to trigger a roaming

In order to roam, a client needs to know about other APs providing access to the same SSID. To do this, the client needs to "scan":

* Active scan: The client changes its radio to a new channel and broadcasts a probe request. It usually waits 10ms for any responses.
  * Directed probe: The probe is sent for a specific SSID
  * Broadcast probe: The probe is sent to null SSID and all APs should respond with the SSIDs they support
* Passive scan: the client changes its radio to a new channle and waits for a periodic beacon. It usually waits for 100ms. Due to the longer wait time most clients prefer Active scanning

During a channel scan the client is unable to transmit or receive data. To reduce the impacts clients can do:

* Background scanning: Scannig happens only when the client is not transmitting or periodically on a single alternate channel to minimze data loss. This way, the client builds knowledge of available APs and can roam faster when needed.
* On-roam scanning: This occurs when roaming is necessary

## AP/Controller perspective

A Mobility Group (MG)is a collection of Mobiliy Controllers (MCs) accross which romaing needs to be supported. An MC can contain up to 24 WLCs. WLCs in a mobility group forward data traffic among the group which enables romaing between controllers and WLC redundancy.

Romaing inside a Mobility Group is done without the need to reauthenticate. If the client roams to an AP in a different WLC but in the same MG, the client datta will be transfered between WLCs.

Up to 3 MGs can be grouped in a Mobility Domain (MD) that supports up to 72 controllers. Clients can roam between controllers in differetn MGs as long as they are in the same MD.

If a client moves from one WLC to another one in the same MD the client needs to reauthenticate, reassociate and to get a new IP.

Controllers in an MG musth share a few parameters:

* Mobility Domain name
* Version
* CAPWAP mode
* ACLs
* WLANs (SSIDs)

WLCs sned mobility control messages between them using UDP 16666 (unencrypted). User data traffic is transmited using EoIP (IP protocol 97) or CAPWAP (UDP 5246) tunnels.

When a client associates and authenticate to an AP, the controller places an entry for the client in its database. This includes:

* MAC and IP address
* Security context and associations
* QoS contecxts
* SSID (WLAN)
* associated AP

## Types of Roaming

### L2 Roaming

L2 roaming occurs when the client moves from one AP to another but remains in the same subnet.&#x20;

* If the client roams from AP1 to AP2 but ends up on the same WLC, then we have **Intracontroller Roaming.** In this case the controller updates the database with the client's new AP.
* If the client roams from AP1 to AP2 but ends up on a different WLC, then we have **Intercontroller Roaming.** In this case the controlles exchanges mobility messages and client data is copied to the new controller. Intercontroller Roaming should remain transparent to the user unless the session timeout is exceeded or the client sends a DHCP Discover. In this case, POP and PoA move from old WLC to the new WLC.

### L3 roaming

L3 roaming occurs when the client moves from one AP to another and doesn't remain in the same subnet. This means the controller changed so it is an **Intercontroller Roaming.** But in this case, instead of moving the client DB to the new controller, the original controller marks the client with an anchor entry in it's own database.  The DB entry is copeid to the new controller  and marked as a foreign entry. The roam remains transparent to the client which gets to keep it's IP Address. This implies both anchro and foreign controller should have similar network access privileges so the client doesn't have connectivity issues after handoff. In this case the POP remains with the original WLC and the PoA moves to the foreign WLC.&#x20;

### Guest Tunneling (Auto-anchor mobility)

In this scenario you have one WLAN (ususally the Guest WLAN) that is tunneled to a predefined set of controllers to restrict clients to a specific subnet.&#x20;


# Wireless Authentication

Deprecated security standards:

* **WEP - Wired Equivalent Privacy** is weak and easily breakable so it is considered deprecated and shouldn't be used.
* **WPA - Wi-FI Protected Access** is also deprecated

Current security standards

* **WPA2** - successor of WPA. It implements the 802.11i security standard&#x20;

## WPA2

### WPA2 Personal mode (PSK)

With Personal mode, WPA uses a pre-shared key (PSK) that needs to be statically configured on client devices.

### WPA2 Enterprise mode (802.1X)

With Enterprise mode, [802.1X](/layer-2-technologies/wireless/wireless-authentication/wpa2-802.1x) and [EAP](/security/eap-101) are used for authentication so each user or device is individually authenticated

### WebAuth

With WebAuth guests can be authenticated in a secure way. WebAuth can also be used for clients that don't support 802.1X or clients that fail 802.1X authentication (as a backup mechanism)

### MAB (MAC Address Bypass)

MAB can be used for devices that don't support authentication that require user interaction. In this case the access to the network is allowed based on  the MAC Address of the device.

## WPS (Wi-Fi Protected Setups)

Some devices support the possibility of distributing stronger keys to the clients that want to connect. These methods usually require an admin to enter a small challenge phrase or to push a button (so it has physical access to the AP). These methods are not recommended as they have different weaknesses so WPS should be disabled.


# WPA2 PSK

PSK authentication uses a symmetric encryption which means that the same key and algorrithm used to encrypt the message is used to decrypt it as well.&#x20;

An 802.11 WLAN client will use Open authentication by default. Open authentication uses no keys and doesn't offer end-to-end security. There is no encryption, per-packet authentication or message integrity check.&#x20;

PSK Authentication requires the key to have been shared with the AP and the client before the authentication process starts. The steps to authenticate using PSK are:

1. The client sends an **Authentication Request** to AP
2. The AP then sends a cleartext **challenge phrase** to the client
3. The client encrypts the phrase with the shared key and sends the **encrypted response** it back to the AP
4. The AP decrypts it with the shared key and checks if it matches the original challenge phrase
5. If the phrases match the AP sends an **Authentication Response** to AP
6. The client sends an **Association Request** to the AP
7. The AP sends an **Association Response** to the client
8. A **virtual port is opened** and the client data is now allowed
9. **Data** exchanged between client and AP will be **encrypted** using the same pre-shared key


# WPA2 802.1X

## 802.1X

The 802.1X mechanism is similar to [802.1X for wired networks](/security/switch-security/802.1x). With Wi-Fi, the 3 roles remain the same:

* **Supplicant**: The wireless client
* **Authenticator**: The AP
* **Authentication server**: The server that manages the access requests (it can run on the same host as the authenticator) - typically RADIUS

In a wireless environment the process is as follows:

1. The client (suplicant) sends an Authentication Request just like for Open Authentication.
2. The AP (authenticator) responds with an Authentication Success
3. The client sends an Association Request
4. The AP responds with an Association Reponse that includes an ID.
5. Even though Asscociation is completed, the virtual port is still not allowed to pass any traffic until the 802.1X authentication completes succesfuly. At this step the client can start the process or the authenticator can request the credentials
6. The AP (authenticator) sends the 802.1X traffic encapsulated to the Authentication server while all other network traffic on that port is dropped. The response from the server is sent to the client (this way the client can also authenticate the server)
7. On a succesful authentication the virtual port will be allowed to pass data.

One key aspect of 802.1x is that it will authenticate each supplicant independently and during the authentication process the server and the client derive an individual key that will be used by the client in this session. Because keys are different for each client and session based it is hard for an intruder to get access to the network.

## PKI and Certificate-Based Authentication

* The Certificate Authority (CA) is used to generate digital certificates for users (clients) and the Authentication servers in order to validate their identities
* Clients requests a user certificate from CA and use it to authenticate themseleves to the server using 802.1X
* Servers request a server certificate from CA or it can use a self-signed certificate when it acts as its own CA
* WLCs that are used as the authentication  servers use preinstalled certi


# IPv4 Addressing

[Backup Interfaces](/ipv4/ipv4-addressing/backup-interfaces)

[GRE Tunnels](/ipv4/ipv4-addressing/tunnel-interfaces/gre-tunnels)

[FHRP 101](/ipv4/ipv4-addressing/fhrp-101)

[DHCP 101](/ipv4/ipv4-addressing/dhcp-101)

[DNS 101](/ipv4/ipv4-addressing/dns-101)

[ARP 101](/ipv4/ipv4-addressing/arp-101)

[IPv4 101](/ipv4/ipv4-addressing/ipv4-101)

[Tunnel Interfaces](/ipv4/ipv4-addressing/tunnel-interfaces)

[BFD - Bidirectional Forwarding Detection](/ipv4/ipv4-addressing/bfd-bidirectional-forwarding-detection)

[NSF - Non Stop Forwarding](/ipv4/ipv4-routing/how-the-routing-table-is-built/nsf-non-stop-forwarding)


# Backup Interfaces

## In Theory

A backup interface is an interface that stays inactive as long as the primary interface is in “up/up” state, but becomes active when the primary interface’s protocol status becomes “down”.\
When the protocol on the primary interface comes back up, the backup interface moves back in a “standby mode”.\
You should know that if the primary interface is administratively down, the backup interface won’t come up. So to test, you have to do something on the other end.\
The backup interface can be configured to come up not only when the primary interface is down, but also when it’s utilization reaches a certain threshold:

```
R(config-if)# backup load {enable_threshold|never} {disable_load|never}
```

To prevent link flapping you can delay the switchover using:

```
R(config-if)# backup delay {enable_delay|never} {disable_delay|never}
```

## In Practice

Here’s an example:

[![Backup Interface Example](https://nyquist.eu/wp-content/uploads/2012/02/BackupInterface.png)](https://nyquist.eu/wp-content/uploads/2012/02/BackupInterface.png)

Backup Interface Example

R1 and R2 are connected over Fa0/0 and Seria1/0 interfaces. We will configure Fa0/0 with ip addresses in 12.0.0.0/24 range S1/0 with addresses in 21.0.0.0/24 range. To test connectivity we will enable one loopback on each router and start rip to to advertise the routes from one to another

```
!On R1:
R1(config)# interface Fa0/0
R1(config-if)# ip address 12.0.0.1 255.255.255.0
R1(config-if)# no shut
R1(config-if)# exit
R1(config)# interface Serial1/0
R1(config-if)# ip address 21.0.0.1 255.255.255.0
! On newer IOS version, no need to specify clock rate on the DCE
R1(config-if)# no shut
R1(config)# interface Lo0
R1(config-if)# ip address 1.1.1.1 255.255.255.255
R1(config-if)# exit
R1(config-if)# router rip
R1(config-router)# network 0.0.0.0
!On R2:
R2(config)# interface Fa0/0
R2(config-if)# ip address 12.0.0.2 255.255.255.0
R2(config-if)# no shut
R2(config-if)# exit
R2(config)# interface Serial1/0
R2(config-if)# ip address 21.0.0.2 255.255.255.0
R2(config-if)# no shut
R2(config)# interface Lo0
R2(config-if)# ip address 2.2.2.2 255.255.255.255
R2(config-if)# exit
R2(config-if)# router rip
R2(config-router)# network 0.0.0.0
```

Shortly, we should be able to ping each router’s loopback interface, from the other one:

```
!On R1:
R1#ping 2.2.2.2
Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 2.2.2.2, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 12/18/24 ms
!On R2:
R2#ping 1.1.1.1
Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 1.1.1.1, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 16/20/24 ms
```

Let’s verify the routing tables, also.

```
On R1:
R1#sh ip route
1.0.0.0/32 is subnetted, 1 subnets
C 1.1.1.1 is directly connected, Loopback0
R 2.0.0.0/8 [120/1] via 21.0.0.2, 00:00:15, Serial1/0
[120/1] via 12.0.0.2, 00:00:04, FastEthernet0/0
21.0.0.0/24 is subnetted, 1 subnets
C 21.0.0.0 is directly connected, Serial1/0
12.0.0.0/24 is subnetted, 1 subnets
C 12.0.0.0 is directly connected, FastEthernet0/0
!On R2:
R2#sh ip route
R 1.0.0.0/8 [120/1] via 21.0.0.1, 00:00:15, Serial1/0
[120/1] via 12.0.0.1, 00:00:06, FastEthernet0/0
2.0.0.0/8 is variably subnetted, 2 subnets, 2 masks
C 2.2.2.2/32 is directly connected, Loopback0
R 2.0.0.0/8 [120/1] via 12.0.0.1, 00:02:27, FastEthernet0/0
21.0.0.0/24 is subnetted, 1 subnets
C 21.0.0.0 is directly connected, Serial1/0
12.0.0.0/8 is variably subnetted, 2 subnets, 2 masks
C 12.0.0.0/24 is directly connected, FastEthernet0/0
R 12.0.0.0/8 [120/1] via 21.0.0.1, 00:03:07, Serial1/0
```

Notice the 2 routes installed for the loopback addresses, one for each physical link.

### The good interface

Now let’s enable Fa0/0 as backup for S1/0 on R1. Let’s also start debugging on R1:

```
R1# debug backup
R1# conf t
R1(config)# interface S1/0
R1(config-if)# backup interface Fa0/0
```

As soon as we set Fa0/0 as the backup interface of S1/0, the backup interface goes down, in standby mode:

```
*Mar  1 01:00:25.015: BACKUP(Serial1/0): changed state to "initializing"
*Mar  1 01:00:25.015: BACKUP(Serial1/0): secondary interface (FastEthernet0/0) configured
*Mar  1 01:00:27.015: BACKUP(Serial1/0): event = timer expired on primary
*Mar  1 01:00:27.019: BACKUP(Serial1/0): secondary interface (FastEthernet0/0) moved to standby
*Mar  1 01:00:27.023: BACKUP(Serial1/0): changed state to "normal operation"
*Mar  1 01:00:29.019: %LINK-5-CHANGED: Interface FastEthernet0/0, changed state to standby mode
*Mar  1 01:00:30.019: %LINEPROTO-5-UPDOWN: Line protocol on Interface FastEthernet0/0, changed state to down
*Mar  1 01:00:30.019: BACKUP(FastEthernet0/0): event = secondary interface went 
```

We can see the status with:

```
R1# show ip interface brief
Interface IP-Address OK? Method Status Protocol
FastEthernet0/0 12.0.0.1 YES manual standby mode down
Serial1/0 21.0.0.1 YES manual up up
Loopback0 1.1.1.1 YES manual up up
```

Now let’s shut down the serial link on R2:

```
R2(config)# interface serial0/0
R2(config-if)# shut
```

Now, based on the keepalive mechanism, the serial link on R1 will move the link into an “up/down” state in about 30 seconds (3 missed keepalives) and will move the backup interface in forwarding mode:

```
*Mar  1 01:08:53.439: %LINEPROTO-5-UPDOWN: Line protocol on Interface Serial1/0, changed state to down
*Mar  1 01:08:53.443: BACKUP(Serial1/0): event = primary interface went down
*Mar  1 01:08:53.443: BACKUP(Serial1/0): changed state to "waiting to backup"
*Mar  1 01:08:53.447: BACKUP(Serial1/0): event = timer expired on primary
*Mar  1 01:08:53.459: BACKUP(Serial1/0): secondary interface (FastEthernet0/0) made active
*Mar  1 01:08:53.459: BACKUP(Serial1/0): changed state to "backup mode"
*Mar  1 01:08:55.447: %LINK-3-UPDOWN: Interface FastEthernet0/0, changed state to up
*Mar  1 01:08:56.447: %LINEPROTO-5-UPDOWN: Line protocol on Interface FastEthernet0/0, changed state to up
*Mar  1 01:08:56.447: BACKUP(FastEthernet0/0): event = secondary interface came up
R1#sh ip int brie
Interface                  IP-Address      OK? Method Status                Protocol
FastEthernet0/0            12.0.0.1        YES manual up                    up
Serial1/0                  21.0.0.1        YES manual up                    down
Loopback0                  1.1.1.1         YES manual up                    up
```

The routing protocol converges and we can ping 2.2.2.2 from 1.1.1.1

```
R1#sh ip route
Gateway of last resort is not set

     1.0.0.0/32 is subnetted, 1 subnets
C       1.1.1.1 is directly connected, Loopback0
R    2.0.0.0/8 [120/1] via 12.0.0.2, 00:00:13, FastEthernet0/0
     12.0.0.0/24 is subnetted, 1 subnets
C       12.0.0.0 is directly connected, FastEthernet0/0
R1#ping 2.2.2.2 source 1.1.1.1

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 2.2.2.2, timeout is 2 seconds:
Packet sent with a source address of 1.1.1.1
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 16/18/24 ms
```

When we bring back up the serial interface on R2, Serial 1/0 will come up on R1 and Fa0/0 will move to standby mode again:

```
R2(config)#int s1/0
R2(config-if)#no shut
!On R1:
*Mar  1 01:18:43.423: %LINEPROTO-5-UPDOWN: Line protocol on Interface Serial1/0, changed state to up
*Mar  1 01:18:43.431: BACKUP(Serial1/0): event = primary interface came up
*Mar  1 01:18:43.431: BACKUP(Serial1/0): changed state to "waiting to revert"
*Mar  1 01:18:43.439: BACKUP(Serial1/0): event = timer expired on primary
*Mar  1 01:18:43.443: BACKUP(Serial1/0): secondary interface (FastEthernet0/0) moved to standby
*Mar  1 01:18:43.443: BACKUP(Serial1/0): changed state to "normal operation"
*Mar  1 01:18:45.443: %LINK-5-CHANGED: Interface FastEthernet0/0, changed state to standby mode
*Mar  1 01:18:46.443: %LINEPROTO-5-UPDOWN: Line protocol on Interface FastEthernet0/0, changed state to down
*Mar  1 01:18:46.443: BACKUP(FastEthernet0/0): event = secondary interface went
R1#sh ip int brie
Interface                  IP-Address      OK? Method Status                Protocol
FastEthernet0/0            12.0.0.1        YES manual standby mode          down
Serial1/0                  21.0.0.1        YES manual up                    up
Loopback0                  1.1.1.1         YES manual up                    up 
```

### The bad interface

Things worked as expected when we set a backup for the serial interface. Now let’s try setting the serial interface as the backup for the ethernet interface:

```
R1(config)# interface serial1/0
R1(config-if)# no backup interface
R1(config-if)# exit
R1(config)# interface Fa0/0
R1(config-if)# backup interface serial1/0
R1(config-if)# end
R1# show ip int brie
Interface                  IP-Address      OK? Method Status                Protocol
FastEthernet0/0            12.0.0.1        YES manual up                    up
Serial1/0                  21.0.0.1        YES manual standby mode          down
Loopback0                  1.1.1.1         YES manual up                    up  
```

Thinks work as expected, now let’s shut the FastEthernet interface on R2:

```
R2(config)# interface fa0/0
R2(config-if)# shut
```

And now we wait…\
When you have waited long enough, you shoud have noticed that the FastEthernet interface on Fa0/0 never went down. The keepalive mechanism on Ethernet links is not used to test connectivity with another host, but to see if the interface can send and receive Ethernet frames. This is because Ethernet links are not considered point-to-point interfaces and they are expected to find more than one neighbor on the link. Since the link will always be up, the backup interface will remain in standby mode and will not be used for forwarding.

The same thing would happen with other Multipoint interfaces, like the Frame Relay physical interface or the multipoing subinterface. A point-to-point subinterface would move the protocol status to down when the DLCI assigned to it si not active.

The solution here is to use a more advanced tracking system, like **Enhanced Object Tracking**


# FHRP 101

## HSRP

HSRP provides a virtual MAC address and a virtual IP address that is shared among a group of routers in order to have a HA infrastructure for the default gateway in a subnet.

### Starting HSRP

```
R(config-if)# standby [GROUP] [ip GROUP-IP [secondary]]
! Default GROUP = 0
! The Group IP can be discovered from other HSRP routers
! the GROUP-IP can't be the interface IP
```

### Timers

```
R(config-if)# standby delay minimum SEC reload SEC
! mimium SEC = time to wait after interface comes up, before HSRP is started
! reload SEC = time to wait after router reloads, before HSRP is started
R(config-if)# standby [GROUP] [msec] HELLO-TIME HOLD-TIME 
! default HELLO-TIME = 3 sec
! default HOLD-TiME = 10 sec
```

Timers are usually learned from the active router. Millisecond timers can only be learned when using version 2. Otherwise, they must be configured on all routers.

### Election process

An election process takes place, where the primary router is elected. Only one device will be elected as the primary router. It will receive and forward packets destined for the Group IP. At the same time, a standby router is elected. It will monitor if the primary router is still reachable and if the active router fails, the standby router takes over and a new router is elected to be the standby router.

The router with the highest priority will be elected as the primary router. The default priority or all routers is 100. In case of a tie, the router with the highest IP Address will becom the primary router.

```
R(config-if)# standby GROUP priority PRI
! Default: 100
```

By default if a router with higher priority comes up, it will not become the primary router unless it is configured for preemption:

```
R(config-if)# standby GROUP preempt [delay {minimum | reload | sync} SEC]
! timers are used to delay the preemption process
```

A standby router with equal priority but a higher IP address will still not preempt the primary router.

### Tracking

HSRP can track an interface or an object and will reduce the priority with a configurable value when the interface or the object state goes down:

```
R(config-if)# standby GROUP track INTERFACE [DECREMENT]
R(config-if)# standby GROUP track OBJECT [decrement DECREMENT]
! Default DECREMENT = 10
```

HSRP can also track other objects using:

### Authentication

HSRP authentication can use clear text or md5 hashes. For md5, you can use ca key-chain or a key-string:

```
R(config)# standby GROUP authentication {text STRING | md5 {key-string STRING| key-chain CHAIN}}
```

With an MD5 key, a hash is computed on a portion of each message and it is sent along with the message. The receiving peer performs the same hash on the received message&#x20;

### Versions

|                             | HSRP v1                                     | HSRP v2                                     |
| --------------------------- | ------------------------------------------- | ------------------------------------------- |
| Sends multicast messages to | 224.0.0.2                                   | 224.0.0.102                                 |
| Supported groups            | 0-255                                       | 0-4095                                      |
| virtual MAC address         | <p>0000.0c07.acXX <br>XX = Group Number</p> | <p>0000.0C9F.FXXX<br>XXX = Group Number</p> |
| Keepalive timers            | Doesn't support msec                        | Supports msec                               |

```
R(config)# standby version {1|2}
```

### HSRP Message

While talking to each other, HSRP enabled rotuers use the following messages:

* **Coup** When a standby router wants to assume the function of the active router, it sends a coup message.
* **Hello** The hello message conveys to other HSRP routers the HSRP priority and state information of the router.
* **Resign** A router that is the active router sends this message when it is about to shut down or when a router that has a higher priority sends a hello or coup message.

### HSRP States

```
R# show standby [brief]
```

* **Active** – The router is performing packet-transfer functions
* **Init or Disabled** – The router is not yet ready or able to participate in HSRP, possibly because the associated interface is not up. HSRP groups configured on other routers on the network that are learned via snooping are displayed as being in the Init state. Locally configured groups with an interface that is down or groups without a specified interface IP address appear in the Init state
* **Learn** – The router has not determined the virtual IP address and has not yet seen an authenticated hello message from the active router. In this state, the router still waits to hear from the active router.
* **Listen** – The router is receiving hello messages.
* **Speak** – The router is sending and receiving hello messages
* **Standby** – The router is prepared to assume packet-transfer functions if the active router fails

## VRRP

VRRP is an open standards implementation that is very similar to HSRP. It uses the terms master/backup instead of primary/standby. VRRP uses a virtual mac address in the format 0000.5E00.01XX, where XX is the group number.\
Most configurations are similar to HSRP, except they start with the vrrp keyword:

```
R(config-if)# vrrp GROUP ip GROUP-IP [secondary]
! the GROUP-IP can be the interface IP
```

VRRP advertisements are sent to 224.0.0.18 with protocol number 112. By default they are sent every 1 second. Default holdtime is of 3 seconds

```
R(config-if)# vrrp GROUP timers advertise [msec] INTERVAL
! The backup routers can learn the advertise interval from the Master Router:
R(config-if)# vrrp GROUP timers learn
```

Cisco devicese allow msec timers for VRRP although this is non-standard.

VRRP can only track objects, not interfaces:

```
R(config-if)# vrrp GROUP track OBJECT [decrement DECREMENT]
```

Also, VRRP is preemptive by default, which is different than HSRP.

## GLBP

GLBP is a Cisco proprietary protocol. The advantage of GLBP is that it additionally provides load balancing over multiple routers (gateways) using a single virtual IP address and multiple virtual MAC addresses. The forwarding load is shared among all routers in a GLBP group rather than being handled by a single router while the other routers stand idle.

An AVG (Active Virtual Gateway) will be elected using the same mechanics as the HSRP primary or the VRRP master router. The difference is that AVG’s role is to maintain a list of maximum 4 AVF (Active Virtual Forwarder) and assignes a MAC address to them, in the format 0007.b4XX.XXYY (XXXX = GLBP Group, YY – VF Number). The AVG will reply to ARP requests with the MAC address assigned to the AVFs, thus achieving load balancing.

A router that is assigned a MAC address will be a primary AVF, while the other routers can take over the MAC address if the primary AVF fails.\
Most GLBP configurations are similar to HSRP and VRRP

```
R(config-if)# glbp GROUP ip [GROUP-IP [secondary]]
```

On the AVG you can set the load-balancing method, using:

```
R(config-if)# glbp GROUP load-balancing [host-dependent | round-robin | weighted]
! round-robin - the next available MAC is used (default)
! weighted - proportionally to the weight
! host-dependent - each client will always receive the same MAC in the ARP reply
```

When using weighted load-balancing you can define a WEIGHT for each router:

```
R(config-if)# glbp GROUP weighting WEIGHT [lower LOWER][upper UPPER]
! A router will stop forwarding when the weight drops under the LOWER threshold
! A router will start forwarding when the weight becomes higher than the UPPER threshold
R(config-if)# glbp GROUP weighting track OBJECT [decrement DECREMENT]
! A tracked object will decrement the weight 
```

Premption is off by default for AVG, but is on by default for AVF, with a delay of 30 seconds.

## GDP

Gateway Discovery Protocol is a feature that enables a host to listen for routing protocol advertisements and select a default gateway. A router must disable ip routing before using GDP:

```
R(config)# no ip routing
R(config)#ip gdp {rip|eigrp|irdp [multicast]}
! For RIP, only v1 is supported
```

To verify, use on the host:

```
R# sh ip route
```

### IRDP

Besides listening to RIP or EIGRP messages, GDP can use IRDP to discover available gateways. In order to configure a router to send IRDP (ICMP Router Discovery Protocol) messages, use:

```
R(config-if)# ip irdp
R(config-if)# ip irdp address IP-ADDR PREFERENCE
```

There is no preemption with IRDP. Instead, the PREFERENCE value is used only to chose between different addresses advertised by the oldest router (if it is configured to send advertise several addresses). A host will choose as default gateway the address with the lowest positive preference, or if they are all negative, the lowest negative.

IRDP messages are sent as broadcasts by default, but the router can be configured to send them as multicast to 224.0.0.1

```
R(config-if)# ip irdp multicast
```

The hosts will chose the IRDP\
To verify, use on the router advertising IRDP:

```
R# sh ip irdp [INTERFACE]
```


# DHCP 101

## DHCP Server

### DHCP Pools

On a router, you have to create one or more pools of DHCP addresses available for lease. When a DHCP server receives a DHCP request, it will know what pool to use based on the IP address of interface that received it. If the request came from a host that is not on that subnet, the router will try to match the GIADDR field which contains the IP address of the interface that received the DHCP request on the relay server.\
To define a pool, use:

```
R(config)# ip dhcp pool POOL
```

Now you can define the pool attributes:

```
R(dhcp-config)# network NETWORK-ADDR [NETMASK]
R(dhcp-config)# domain-name DOMAIN
R(dhcp-config)# dns-server SERVER1 [SERVER2 ...]
R(dhcp-config)# default-router GATEWAY1 [GATEWAY2 ...]
R(dhcp-config)# option CODE [instance NUMBER] {ascii STRING | hex STRING | IP-ADDR}
R(dhcp-config)# lease {DAYS [HOUR [MINUTES]]|infinte}
! default - 1 DAY
R(dhcp-config)# update {arp|dns}
! updates local arp and dns info with data from dhcp server.
! update arp is used for Authorized ARP - See ARP 101 article
```

You can exclude some IP’s from being used in the DHCP pool using:

```
R(config)# ip dhcp excluded-address START-IP [END-IP]
```

### Manual Bindings

To make manual bindings you have to create additional pools, one for each static binding. In these pools you must specify the host IP-Address and an identifier for the host – Client ID or Hardware Address. Clients are matched by the hardware address only if they don’t send a client ID. Cisco routers always send a Client ID!

```
R(config)# ip dhcp pool STATIC-HOST-1
R(dhcp-config)# host IP-ADDR [NETMASK]
R(dhcp-config)# client-identifier CLIENT-ID
R(dhcp-config)# hardware-address MAC-ADDRESS [PROTOCOL-TYPE|HARDWARE-NUMBER]
R(dhcp-config)# client-name HOSTNAME
```

Defining one POOL for each client can become an administrative nightmare. Another option is to use static mappings from a file similar to the one that the router saves when using the database agent.\
To enable the use of a static file, use this command in a Network Pool:

```
R(dhcp-config)# origin file URL
```

### ODAP – On Demand Address Pool

A pool can be configure to use one or more subnets from another DHCP or AAA server. This is usually used by providers to assign addresses from the same pool to different customers. To define an ODAP pool, use:

```
R(config)# ip dhcp pool ODAP-POOL
R(dhcp-config)# origin {dhcp|aaa|ipcp}
! Optionally the pool can be assigned to a vrf
R(dhcp-config)# vrf VRF-NAME
! Now you can import the dhcp options from the server into the current pool
R(dhcp-config)# import all
```

The ODAP server must also be configured to reply with the requested subnets. To do this, configure the server pool with:

```
R(config)# ip dhcp pool SERVER-POOL
R(dhcp-config)# network NETWORK-ADDR NETMASK
R(dchp-config)# subnet prefix LEN
! Defines size of the subnet that can be allocated to ODAP clients
```

### DHCP Relay Agent

By default, the list of DHCP bindings is kept in memory by the router and they are lost once the router reloads. You can enable the router to use a database agent, that is a location where the router can save the list of the bindings. This file can also be used to recover the list in case of reloads. To enable this feature, use:

```
R(config)# ip dhcp database URL
! URL can be a local or a remoate location. Ex: flash:dhcp.txt
```

### DHCP Classes for Option 82

When a host requests an address via DHCP, the server uses the incoming interface IP or the GIADDR field in the DHCP request (filled by a relay agent with it’s incoming interface IP) that would match the network address defined for a pool. Additionally, it can use the client-id or hardware-address for specific manual bindings.\
Option 82 is a special field in a DHCP packet that can contain additional information that may identify a host. Switches can be configured to insert information in this field (and in the Relay Information Option field). See [DHCP Snooping](https://nyquist.eu/dhcp-snooping-and-dai/). To process this information, the DHCP server on a router needs to use DHCP classes. By default, DHCP classes are on, but if disabled, they can be enabled again with:

```
R(config)# ip dhcp use class
```

When you configure a class, you actually define the option fields that the router will compare when it receives a DHCP request. The match is done bitwise.

```
R(config)# ip dhcp class CLASS-NAME
! Define option fields
R(config-dhcp-class)# option OPTION-NUMBER hex OPTION-HEX-VALUE [mask BITMASK]
! Add Relay Agent Information fields:
R(config-dhcp-class)# relay agent information
R(config-dhcp-class-relayinfo)# relay-information hex RELAY-HEX-VALUE [mask BITMASK]
! If missing, it matches anything
```

Then apply the class to an existing DHCP pool. When you do this, the pool will be used for allocation only if at least one class is matched by a DHCP request.

```
R(dhcp-config)# class CLASS-NAME
! You can assign a sub-range of the pool to this class:
R(config-dhcp-pool-class)# address range START END
! Or forward the requests to another DHCP server:
R(config-dhcp-pool-class)# relay target OTHER-DHCP-SERVER
```

The value of the Option82 field is also added back to the DHCP reply messages that the server sends to its clients.

If the relay agent inserts Option 82 but doesn’t add GIADDR field, the router will drop the DHCP message unless you configure it to trust such messages:

```
! Globally, for all interfaces:
R(config)# ip dhcp relay information trust-all
! or just for one interface:
R(config-if)# ip dhcp relay information trusted
```

Another option is to disable the insertion of Option82 on the switch:

```
Sw(config)# no ip dhcp snooping information option 
```

## DHCP Relay Agent

You can forward DHCP requests to another server using the feature of [forwarding UDP protocols](https://nyquist.eu/multicast-101/#61_Convert_broadcast_to_unicast_8211_Helper_Address). For DHCP, the following command should be enough:

```
R(config-if)#ip helper-address REMOTE-SERVER
```

The requests will be sent as unicast and the relay agent will add the GIADDR information (IP address of the interface that received the DHCP request)\
Another option is to make the router forward DHCP requests from within a pool:

```
R(config)# ip dhcp pool POOL-NAME
! The relay source will match be used to match the incoming interface or GIADDR
R(dchp-config)# relay source NETWORK-ADDR MASK
! The relay destination will define where the DCHP request is forwarded
R(dhcp-config)# relay destination REMOTE-SERVER
```

When using classes, you can define a relay target for each class:

```
R(config-dhcp-pool-class)# relay target OTHER-DHCP-SERVER
```

### Option 82

The DHCP relay can be configured to add Option82 information to the DHCP requests that it forwards. You can enable this globally or per interface:

```
R(config)# ip dhcp information option
Rack1R4(config-if)#ip dhcp relay information option-insert [none]
```

Also by default, the router will also check the reply messages from the server before forwarding them to the host. If they don’t have the Option82 information echoed back in the reply packet, it will be dropped. You can disable this check globally, or per interface:

```
R(config)# no ip dhcp relay information check
R(config-if)# ip dhcp relay information check-reply [none]
! Use none to disable
```

A Cisco switch can insert Option82 information into a DHCP request. When it receives DHCP requests that already have this information attached, the router will replace it with its own. You can configure how it should treat these packets with any of the following commands:

```
R(config)# ip dhcp relay information policy {drop|keep|replace}
R(config-if)# ip dhcp relay information policy-action {drop|keep|replace}
```

Again, if the relay agent inserts Option 82 but doesn’t add GIADDR field, the router will drop the DHCP message unless you configure it to trust such messages:

```
! Globally, for all interfaces:
R(config)# ip dhcp relay information trust-all
! or just for one interface:
R(config-if)# ip dhcp relay information trusted
```

### 3DHCP Client

Settings for DHCP clients can be configured globally, or on each interface:

```
! Global DHCP Client config
R(config)# ip dhcp-client ?
  broadcast-flag     Set the broadcast flag
  default-router     Set DHCP default router related information
  forcerenew         Enable forcerenew client processing
  network-discovery  Configure parameters for network discovery
  update             Configure automatic updates
```

```
! Interface DHCP Client config
R(config-if)# ip dhcp client ?
  class-id   Specify Class-ID to use
  client-id  Specify Client-ID to use
  hostname   Specify hostname to use
  lease      Requested address lease time
  mobile     Mobile client configuration parameters
  request    Specify options (not) to request
  route      Options for routes installed by dhcp
  update     Dynamically update information
```

Another option is to specify the information when you enable the DHCP client:

```
R(config-if)# ip address dhcp [client-id INTERFACE][hostname NAME]
```

A DHCP router will always include a client-id in it’s request. This means that a Cisco DHCP server will not use it’s mac-address when searching for a host pool used for address assignment.\
You can see what is the value of the client-id sent by the router in it’s DHCP requests with:

```
R#sh dhcp lease
```

If you want to match this value on the DHCP server, you will have to use the hex-value but in a dotted format. So a better option would be to change the client id before sending the request.\
To see debug information for the client, use:

```
R# debug dhcp
```

You can use 2 exec commands to release or renew the DHCP address:

```
R# release dhcp INTERFACE
R# renew dhcp INTERFACE
```


# DNS 101

## Defining hosts and domains locally

To define a static host name to address mapping, use the following command:

```
R(config)# ip host NAME [TELNET-PORT] ADDRESS
```

For hosts that are accessed without a domain name at the end, (hostname.domain-name), you can define a default domain name or a list of domain names to be used, using one of the following commands:

```
R(config)# ip domain name DOMAIN
R(config)# ip domain list DOMAIN
! Domain list is preferred over domain name. 
```

## Using a DNS Server

To define a DNS server and make the router work as a DNS client, use:

```
R(config)# ip name-server SERVER1 [SERVER2 ...]
```

The local mappings will be used first, and if there is no match, the name-server will be queried.

For DNS lookups you can define:

```
R(config)# ip domain timeout SEC
R(config)# ip domain retry NUMBER
```

By default, if there are multiple DNS servers defined, the first one will be used by default, while the other servers will only be used in case of failure of the previous servers.\
For one host a router can have multiple IP addresses that it is resolved to. By default, only the first one is used. You can use each IP address in a round-robin fashion if you enable

```
R(config)# ip domain round-robin
```

To see the current dns cache, use:

```
R# show ip hosts
```

Of course, lookups can be disabled altogether, using:

```
R(config)# no ip domain-lookup
```

## Making the router a DNS Server

To configure the router as a server, use:

```
R(config)# ip dns server
```

The router will respond to DNS requests with data from its statically configured hosts or from the responses cached from the other DNS servers.

DNS Spoofing is an option that is enabled only if the domain lookups are disabled, if no name servers are configured, or if there is no route to them. The router will respond to all DNS requests with the configured IP or with the interface IP:

```
R(config)# ip dns spoofing [IP-ADDRESS]
```


# ARP 101

## ARP

ARP is a protocol used on broadcast networks such as Ethernet, Token Ring or FDDI that is used to map L3 Addresses (like IP) to layer 2 Addresses (like Ethernet MAC).\
When a host needs to send traffic to another host, it knows its Layer 3 address, but it needs to find out the Layer 2 address in order to encapsulate the frame. In order to find the Layer 2 Address, the host will send an ARP Request to the Layer 2 asking “Who has the IP Address x.x.x.x?”. All hosts in the broadcast domain will receive this message, but only the one that was assigned that specific L3 address should respond with a unicast ARP Reply.

When a host receive an ARP message, either ARP Request(broadcast) or ARP Reply(unicast) it updates its ARP Cache. This cache contains all the L3 to L2 mappings that the host knows about. When it needs to send a packet, the host will look in the ARP Cache to find the appropriate mapping. If there is no mapping for the destination IP Address it will send an ARP Request. If it doesn’t receive an ARP Reply, then L2 encapsulation will fail. Entries in the ARP Cache can be dynamic (with a limited lifetime) or static (permanent).\
To define static entries, use:

```
R(config)# arp IP-ADDRESS MAC ENCAPSULATION-TYPE ...
! For Ethernet, ENCAPSULATION-TYPE = arpa
```

To configure the timeout of dynamic ARP entries, use:

```
R(config)#arp timeout SEC
!default: 14400 sec = 4 hours
```

You can clear dynamic ARP entires using one of the following commands:

```
! per interface
R# clear arp interface INTERFACE
! all dynamic entires
R# clear arp-cache
```

To see the ARP cahce, use:

```
!All L3 protocols:
R# show arp
! Only IP
R# show ip arp
```

### Authorized ARP

Authorized ARP disables the dynamic update of the ARP cache on an interface. This means that clients connecting to that interface will not be able to communicate with the router unless their MAC address was added to the cache by an authorized process. Authorized processes are static ARP entries and DHCP generated entries.\
To enable DHCP to update the arp-cache, use the following command inside the DHCP Pool:

```
R(dhcp-config)# update arp
```

See [DHCP 101](https://nyquist.eu/dhcp-101/#11_DHCP_Pools) for details.

## Inverse ARP

Inverse ARP is used in Frame Relay or ATM networks and it is used to find the IP Address of the device connected at the other end of a Virtual Circuit. See [Frame Relay](https://nyquist.eu/frame-relay-101-part-2/)

## Reverse ARP

Reverse ARP is used when a host doesn’t know its IP Address and works similar to DHCP. The host will send a RARP messages with its MAC Address and expects to receive a reply from a RARP server letting it know what IP Address should use.

## Proxy ARP

Proxy ARP is used when a host needs to communicate with another host that is not in the same broadcast domain. A router can detect that the destination IP is not in the same broadcast domain and if it has a route to that destination it can respond with its own L2 address. The packets for the L3 destination will reach the router which will decapsulate and reencapsulate them before sending them over another interface.\
Proxy ARP is enabled by default on routed interfaces. You can disable proxy ARP per interface, using:

```
R(config-if)#no ip proxy-arp
```

or globally:

```
R(config)# ip arp proxy disable
```

### Local Proxy ARP

Even when configured with proxy-apr, a router will not respond to ARP requests for destination that are on the same incoming interface. However, in some situations (like when having a router connect hosts in an isolated Private VLAN) you might need the router to respond to such ARP requests. To enable this, use:

```
R(config-if)# ip local-proxy-arp
```


# IPv4 101

## Setting an IP Address

```
R(config)#ip address IP-ADDR NETMASK
```

Any combination of IP-ADDR and NETMASK can be used as long as the **host** portion of the address is not all zeros. One exception is allowed, when using a /31 mask.\
By default, the router will accept combinations that result in a **subnet** portion of the address with all zeros, but this behavior can be disabled using:

```
R(config)#no ip subnet-zero
```

### Secondary IP Addresses

Multiple IP Addresses can be configured on an interface using the **secondary** keyword:

```
R(config-if)# ip address IP-ADDR NETMASK secondary
```

### IP Unnumbered

This feature allows a router to use the IP address of another interface instead of using a new IP Address. It can only be configured on point-to-point interfaces

```
R(config-if)# ip unnumbered INTERFACE
```


# Tunnel Interfaces

## Tunnel Modes

A tunnel makes two distant devices appear directly connected over a logical interface. When a packet is sent out on the tunnel interface, it is encapsulated in the “carrier” protocol and sent over a physical interface.The most used carrier protocols are GRE, IP-in-IP and IPv6, and this can be set using the tunnel mode.

When configuring a tunnel you must set a tunnel source, a destination and the carrier protocol:

```
R(config)# interface TUNNEL
R(config-if)# tunnel source {SRC-ADDR|SRC-INTERFACE}
R(config-if)# tunnel destination DEST-ADDR
R(config-if)# tunnel mode MODE
```

The tunnel source can be defined as either the Layer3 address or as an interface, but the destination can only be an address on the remote device.\
The tunnel mode specifies the carrier protocol. For IPv4, the most commonly used methods of tunneling are GRE and IPIP.

```
R(config-if)# tunnel mode {gre ip|ipip}
! gre ip - Encapsulates IPv4 in GRE
! ipip - Encapsulates IPv4 in IPv4
```

After this, you can define the encapsulating protocol’s address on the tunnel interface:

```
R(config-if)# ip address {unnumbered INTERFACE| IP-ADDR NETMSK}
```

When you have isolated networks running IPv6, you can connect them over an IPv4 backbone using tunnels. Common options are:

```
R(config-if)# tunnel mode {gre ipv6|ipv6ip [6to4|auto-tunnel|isatap]}
! gre ipv6 - Encapsulates IPv6 in GRE
! ipv6ip - Encapsulates IPv6 in IPv4
```

More details about configuring IPv6 tunnels can be found [here](https://nyquist.eu/interconnecting-ipv6-and-ipv4/#2_Tunnels)

Another option is to have ipv6 as the transport protocol:

```
R(config-if)# tunnel mode ipv6
```

The transported protocols can be IPv4 or IPv6, based on the type of address defined on the tunnel interface.

### GRE (Generic Routing Encapsulation)

GRE is defined as IP Protcol 47. It adds a 20 byte IP header and 4 byte GRE header to an existing packet so that it can be routed based on this new information.

#### GRE Keepalives

GRE supports sending and monitoring keepalives to determine the status of a tunnel interface.

```
R(config-if)# keepalive [PERIOD [RETRIES]]
! Default PERIOD: 10 sec
! Default RETRIES: 5
```

## Path MTU Discovery

Tunneling packets means an extra encapsulation header that is added to the packet which can make the packet too big on some links.\
To set the MTU value on a tunnel you can set it manually or use auto-discovery:

```
R(config-if)# ip mtu MTU
R(config-if)# tunnel path-mtu-discovery [age-timer {TIME|infinte}| min-mtu SIZE]
! Default TIME: 10 min
```

The discovery will run for a limited ammount of time unless the infinte keyword is used. Using SIZE, you can configure a minimum size of the discovered MTU that can be accepted. If the timer expires, the router will choose a MTU equal to the default interface MTU-20 Bytes for IP-in-IP or default interface MTU-24 Bytes for GRE.\
Path MTU discovery is only available on GRE and IPIP tunnels.

## VRF Support

The transport protocol of one tunnel can run in one VRF, while the transported protocol can run in another VRF. By default, both the transport and the transported protocols run in the global VRF. You can change the VRF of the transporting protocol with:

```
R(config-if)# tunnel vrf VRF
```

Source and destination addresses of the tunnel must run in this VRF.\
To change the VRF of the transported protocol, use:

```
R(config-if)# ip vrf forwarding vRF
```

The addresses defined on the tunnel will run in this VRF.


# GRE Tunnels

## Configuring GRE

GRE tunnels appear as directly connected logical interfaces to the router even though the traffic that goes into the tunnel will actually be carried over other physical interfaces to the destination. To configure a GRE tunnel you have to set the source and the destination of the tunnel on each end. The source can be an interface or the IPv4 or IPv6 address of an interface. The tunnel destination can only be an IPv4 or IPv6 address on the other router.

```
R(config)# interface TUNNEL-INTERFACE-ID
R(config-if)# tunnel mode gre {ip|ipv6|multipoint}
R(config-if)# tunnel source {INTERFACE|IPV4-ADDRESS|IPV6-ADDRESS/MASKLEN}
R(config-if)# tunnel destination{IPV4-ADDRESS|IPV6-ADDRESS}
```

As you can see, the tunnel can run over IPv4 or IPv6 depending on the tunnel mode. By default a tunnel interface is set to **gre ip**. The multipoint option will be discussed in another article.\
GRE tunnels can also transport IPv4 and IPv6, depending on the Layer 3 addresses that are configured on the tunnel interface:

```
R(config-if)# ip address {IPV4_ADDRESS MASK| IPV6_ADDRESS/MASKLEN}
```

A tunnel interface will come up on one router regardless of the end-to-end connectivity. To have a more reliable status of the interface we should enable keepalives, which are disabled by default. The keepalive mechanism is pretty ingenious and can be enabled on only one side, or with different options on both sides. Read [this article](https://www.cisco.com/en/US/tech/tk827/tk369/technologies_tech_note09186a008040a17c.shtml) for more details.

```
R(config-if)# keepalive [PERIOD [RETRIES]]
```

A limited security mechanism can be implemented, in which the routers must attach a preshared numerical key to each packet. If the keys don’t match, the destination will drop the packets.

```
R(config-if)# tunnel key KEY-VALUE
```

Normally, when bringing up a tunnel between 2 routers we create a loop because there are now 2 logical paths for the traffic: One over the physical links and one over the logical tunnel interface. One thing that is very probable to happen is to have a “**recursive routing**” loop. This will happen if the routing mechanism decides to use the logical tunnel interface to reach the defined tunnel destination. Cisco IOS detects this and will move the tunnel interface to a down state. This will probably disable the route through the tunnel interface which will re-enable the Tunnel interface and then the route through the interface will come up and IOS will again detect the “recursive routing” situation, and so on.

## The example

![GRE Example Topology](/files/61ZkSpNQe46mPIEJHYQH)

Each router will be configured with a loopback address. We will create a tunnel between R1 and R3 and we will use EIGRP for route distribution.\
Here’s the starting configs that only include IP addressing information

```
! On R1:
interface Loopback0
 ip address 1.1.1.1 255.255.255.255
!
interface FastEthernet0/0
 ip address 12.0.0.1 255.255.255.0
 duplex auto
 speed auto
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
! On R2
interface Loopback0
 ip address 2.2.2.2 255.255.255.255
!
interface FastEthernet0/0
 ip address 12.0.0.2 255.255.255.0
 duplex auto
 speed auto
!
interface FastEthernet0/1
 ip address 23.0.0.2 255.255.255.0
 duplex auto
 speed auto
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
! On R3:
interface Loopback0
 ip address 3.3.3.3 255.255.255.255
!
interface FastEthernet0/1
 ip address 23.0.0.3 255.255.255.0
 duplex auto
 speed auto
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
```

Now let’s create the tunnel interfaces on both ends:

```
!On R1:
R1(config)# interface Tunnel0
R1(config-if)# tunnel source 12.0.0.1
R1(config-if)# tunnel destination 23.0.0.3
R1(config-if)# ip address 13.0.0.1 255.255.255.0
!On R3:
R3(config)# interface Tunnel0
R3(config-if)# tunnel source 23.0.0.3
R3(config-if)# tunnel destination 12.0.0.1
R3(config-if)# ip address 13.0.0.3 255.255.255.0
```

Let’s test connectivity:

```
! On R1
R1#ping 13.0.0.3

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 13.0.0.3, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 40/56/84 ms
```

OK – the tunnel seems to work – Let’s see how a traceroute looks:

```
R1#traceroute 13.0.0.3 numeric

Type escape sequence to abort.
Tracing the route to 13.0.0.3

  1 13.0.0.3 72 msec *  40 msec
```

As expected, R2 does not show up in the traceroute because the tunnel looks like a directly connected interface. Let’s see how can we reach R3’s loopback:

```
R1#tracer 3.3.3.3 numeric

Type escape sequence to abort.
Tracing the route to 3.3.3.3

  1 12.0.0.2 40 msec 28 msec 20 msec
  2 23.0.0.3 60 msec *  60 msec
```

Hmm, it seams we can actually see R2 in the trace. How come? Let’s look at the routing table:

```
R1#sh ip route
Gateway of last resort is not set

     1.0.0.0/32 is subnetted, 1 subnets
C       1.1.1.1 is directly connected, Loopback0
     2.0.0.0/32 is subnetted, 1 subnets
D       2.2.2.2 [90/409600] via 12.0.0.2, 00:13:14, FastEthernet0/0
     3.0.0.0/32 is subnetted, 1 subnets
D       3.3.3.3 [90/435200] via 12.0.0.2, 00:13:14, FastEthernet0/0
     23.0.0.0/24 is subnetted, 1 subnets
D       23.0.0.0 [90/307200] via 12.0.0.2, 00:13:14, FastEthernet0/0
     12.0.0.0/24 is subnetted, 1 subnets
C       12.0.0.0 is directly connected, FastEthernet0/0
     13.0.0.0/24 is subnetted, 1 subnets
C       13.0.0.0 is directly connected, Tunnel0
```

We see that in order to reach 3.3.3.3 we are using the link to R2, not through R3. Here’s why:

```
R1#sh ip eigrp topology
IP-EIGRP Topology Table for AS(100)/ID(1.1.1.1)

Codes: P - Passive, A - Active, U - Update, Q - Query, R - Reply,
       r - reply Status, s - sia Status

P 3.3.3.3/32, 1 successors, FD is 435200
        via 12.0.0.2 (435200/409600), FastEthernet0/0
        via 13.0.0.3 (297372416/128256), Tunnel0
! Rest of the output omitted
```

According to the EIGRP topology table, we have 2 routes towards 3.3.3.3, one over Fa0/0 through R2, and one over the Tunnel0 interface, through R3, but the metric via the Tunnel0 interface is a lot worse than via Fa0/0 and here’s the explanation:

```
R1#sh int fa0/0 | i BW
  MTU 1500 bytes, BW 10000 Kbit, DLY 1000 usec,
R1#sh int tun0 | i BW
  MTU 1514 bytes, BW 9 Kbit, DLY 500000 usec,
```

The default values for bandwidth and delay on a tunnel interface are much worse than on the physical interface. Let’s change this values to see if we can force the use of the other link:

```
R1(config-if)#delay 100
R1(config-if)#bandwidth 10000
```

Now, it won’t take much until we see these errors:

```
*Mar 1 08:12:36.197: %LINEPROTO-5-UPDOWN: Line protocol on Interface Tunnel0, changed state to up
*Mar 1 08:12:36.237: %DUAL-5-NBRCHANGE: IP-EIGRP(0) 100: Neighbor 13.0.0.3 (Tunnel0) is up: new adjacency
*Mar 1 08:12:45.197: %TUN-5-RECURDOWN: Tunnel0 temporarily disabled due to recursive routing
*Mar 1 08:12:46.197: %LINEPROTO-5-UPDOWN: Line protocol on Interface Tunnel0, changed state to down
*Mar 1 08:12:46.249: %DUAL-5-NBRCHANGE: IP-EIGRP(0) 100: Neighbor 13.0.0.3 (Tunnel0) is down: interface down
```

We have hit the Recursive Routing scenario, in which the router uses the Tunnel interface to reach the Tunnel destination. This is, of course an illegal operation. The Tunnel interface will change state do down. When the routes over the Tunnel interface expire, the router will try to bring up the tunnel interface. The tunnel will come up, a new EIGRP adjacency will form over the Tunnel interface, it will have a better path towards the Tunnel destination and we entered the loop. So, remember that the low bandwidth and high delay values on the Tunnel interface are just a protection mechanism against those kind of loops. OSPF also uses the bandwidth in the default cost formula so we should be safe most of the time (but not always!), but what about RIP? RIP will definitely cause some problems since the default metric only uses hop count, and we will have just one hop between the 2 ends of a tunnel.\
So how can we send traffic over the tunnel interface? One option is to use static routes to force the traffic over the Tunnel interface, but using static routes is not a good idea in large environments. Another option is to Policy Route traffic towards 3.3.3.3 via 13.0.0.3 but is just as unscalable as the static route and it is also more difficult to read out in the config.\
A third option is to use another routing protocol with a lower Administrative Distance than the one running over the Tunnel interface, to reach the tunnel destination.\
Probably the best option is to limit the routes advertised or received over the Tunnel interface, so that we don’t use the Tunnel interface to reach the Tunnel destination. To do this, we should use a distribute-list:

```
R1(config)# ip prefix-list LIST0 deny 12.0.0.0/24
R1(config)# ip prefix-list LIST0 deny 23.0.0.0/24
R1(config)# ip prefix-list LIST0 permit 0.0.0.0/0 le 32
R1(config)# router eigrp 100
R1(config-router)# distribute-list prefix LIST0 in Tun0
R1(config-router)# distribute-list prefix LIST0 out Tun0
```

I filtered on R1 both incoming and outgoing routes over Tunnel0 interface. Another method would have been to filter one way on each router, but the result is the same. Let’s see how that worked out:

```
R1#sh ip route
Gateway of last resort is not set

     1.0.0.0/32 is subnetted, 1 subnets
C       1.1.1.1 is directly connected, Loopback0
     2.0.0.0/32 is subnetted, 1 subnets
D       2.2.2.2 [90/409600] via 12.0.0.2, 00:10:11, FastEthernet0/0
     3.0.0.0/32 is subnetted, 1 subnets
D       3.3.3.3 [90/409600] via 13.0.0.3, 00:09:53, Tunnel0
     23.0.0.0/24 is subnetted, 1 subnets
D       23.0.0.0 [90/307200] via 12.0.0.2, 00:13:45, FastEthernet0/0
     12.0.0.0/24 is subnetted, 1 subnets
C       12.0.0.0 is directly connected, FastEthernet0/0
     13.0.0.0/24 is subnetted, 1 subnets
C       13.0.0.0 is directly connected, Tunnel0
R1#traceroute 3.3.3.3 source 1.1.1.1 numeric

Type escape sequence to abort.
Tracing the route to 3.3.3.3

  1 13.0.0.3 80 msec *  72 msec
```

It finally looks good.

## Conclusions on Recursive Routing

You saw that EIGRP prevents recursive routing situations in small scenarios like this one by using a default metric that is worse through the tunnel than through the physical path. The same happens with OSPF as long as the physical path has a bandwidth higher than 9kbps (default for GRE tunnels). However with RIP, things won’t go as smoothly because it’s metric is based on hop count and the tunnel interface will always count as 1 hop, while the physical path will probably be more than that. So, pay attention when using RIP over GRE tunnels, as it will enter the recursive routing state by default.


# BFD – Bidirectional Forwarding Detection

## What is BFD?

BFD stands for Bidirectional Forwarding Detection and it’s a protocol that is used for rapid detection of link failures when the line-protocol is still “up”. BFD is enabled on interface and creates a BFD session with the neighboring router (BFD Peer). Routing protocols such as EIGRP, OSPF and BGP support BFD and can rapidly detect a link failure without waiting for hold-timers to expire.

## Enabling BFD

To enable BFD per interface, use:

```
R(config)# interface INTERFACE
R(config-if)# bfd interval TX-MSEC min_rx RX-MSEC multiplier COUNT
! TX-MSEC = rate at which BFD packets are sent
! RX-MSEC = Min rate at which BFD packets are expected
! COUNT = how many packets can be missed before considering the link down
```

To verify BFD, use:

```
R# show bfd neighbors [detail]
```

## BFD Support for routing protocols

### EIGRP

To enable BFD support for EIGRP, use:

```
R(config)# router eigrp AS-NUMBER
R(config-router)# bfd {interface INTERFACE|all-interfaces}
```

To verify, use:

```
R# show ip eigrp interfaces ... [details] 
```

### OSPF

To enable BFD support for OSPF, there are 2 options:\
Option 1 is to enable OSPF support for BFD, one-by-one on each interface:

```
R(config)# interface INTERFACE
R(config-if)# ip ospf bfd
```

Option 2 is to enable OSPF support for all BFD interfaces and the disable it on unneeded interfaces:

```
R(config)# router ospf PROC-ID
R(config-router)# bfd all-interfaces
! Will enable BFD support for all interfaces
```

You can disable then OSPF support for BFD, per-interface, with:

```
R(config-if)# ip ospf bfd disable
```

To verify, use:

```
R# show ip ospf [details]
```

#### 2.3BGP

To enable BFD support for BGP, use:

```
R(config)# router bgp AS-NUMBER
R(config-router)# neighbor NEIGH-IP fall-over bfd
```

To verify, use:

```
R# show ip bgp neighbors
```


# IPv4 Routing


# How the routing table is built

## The Routing Process

The routing mechanism involves three different processes: the routing protocols that build the routing table, the routing process that looks up the routing table and the forwarding process that inquires the routing process what to do with a packet.

Several routing protocols can run on one router. Together with the static routes, they will build the routing table. Each protocol will have its own list of destinations and how to reach them. Each protocol will try to add its own destinations to the routing table, but only some of them will make it into the routing table.

A destination is represented by a network number and a network mask. Two destinations are considered different if either the network number or the network mask are different, even if they are part of the same major network. Here’s an example: 10.0.0.0/24 is different than 10.0.1.0/24, but also 192.168.10.0/24 is different than 192.168.10.0/25.

## Building the routing table

When installing a route in the routing table, the following algorithm is used:

1. If there is just one route for a destination, that route is added to the routing table.
2. If there are several routes for a destination, the routes with the lowest administrative distance could be added to the routing table.
3. If there are several routes for a destination and they all have the same administrative distance, then the route with the best metric could be added to the routing table.
4. If there are one or more routes for a destination with equal administrative value and metric, then all these routes are added to the routing table.

| Route Source       | AD  |
| ------------------ | --- |
| Directly Connected | 0   |
| Static Route       | 1   |
| EIGRP summary      | 5   |
| BGP External       | 20  |
| EIGRP Internal     | 90  |
| IGRP               | 100 |
| OSPF               | 110 |
| IS-IS              | 115 |
| RIP                | 120 |
| EGP                | 140 |
| ODR                | 160 |
| EIGRP External     | 170 |
| BGP Internal       | 200 |
| Unknown            | 255 |

### What is the Administrative Distance?

The Administrative Distance (AD) differentiates the routes based on the routing protocol that knows about them. Usually, the default values are used, but they can be changed. The route with the lowest AD could be added to the routing table. The next tiebreaker is the metric.

### What is the metric?

Within the routes that come from the same routing protocol, the metric is the tiebreaker used to differentiate them. Since the metric is calculated differently for each routing protocol, comparing metrics for routes with different AD makes no sense.

## Looking up the best route in the routing table

### Classless routing

With classless routing, a routing table lookup always returns the longest match. This actually means any supernet that contains our destination address. If there is no destination that matches our destination IP, the default route could be used. The default route is a usually defined as 0.0.0.0/0 or 0.0.0.0 0.0.0.0 and is the lowest possible entry in the routing table that will match any route.\
Classless routing is the default behavior of the routing table, but if it was changed, you can use it again, with:

```
R(config)# ip classless
```

### Classful routing

If using classful routing, when a packet arrives, the destination address is used to route the packet.

* Is the major network of the destination address in the routing table?
  * NO: Is there a default route?
    * YES: Use the default route
    * NO: Drop the packet
  * YES: Is there a subnet of the major network, where the destination address fits?
    * NO: Drop the packet
    * YES: Route the packet

To enable classful routing, use:

```
R(config)# no ip classless
```

## Finding the exit interface

Usually, a route for a destination points to the address of the next-hop. The router will have to make another routing table lookup to find out how to route towards the next-hop. This is done recursively until it finds a match pointing to a directly connected network, so it finds the outgoing interface.

In order to send the packet out the outgoing interface, the router must now build the Layer 2 packet. If the outgoing interface is a point-to-point interface (like a HDLC, PPP or frame-relay point-to-point subinterface), the router should have all information to send the packet to the other end.\
If the outgoing interface is a multipoint interface (ethernet, frame-relay physical interface or frame-relay multipoint subinterface) then the router must find the Layer 2 address to use in order to get to the destination.

The layer 2 address to use could be statically defined or it can be discovered using ARP on Ethernet or Inverse ARP on Frame Relay, but there can still be some issues.

Static routes can be configured to point to a next-hop address or to an outgoing interface. If the destination is not on the subnet of the outgoing interface, normal ARP will fail because no host will reply to the ARP requests. In such situations, Proxy ARP can help. A router on the subnet configured for Proxy ARP could respond with its Layer 2 address in order to receive the packet and route it towards the destination. Proxy ARP is enabled by default on Cisco Routers, but depending on the IOS version, it can be enabled or disabled per interface:

```
R(config)# R(config-if)# [no] ip proxy-arp
! The interface must be a Layer 3 interface.
```

## Load sharing

We know that a router will add into its routing table only the best route to the destination \[AD/metric]. But what if there are multiple paths that have the same \[AD/metric]. Which one should be used? It turns out the router can add multiple routes for the same destination, as long as they have the same metric. The number of routes that can exist in the routing table at one time for the same destination depends but it can be changed with:

```
R(config)# maximum-paths VAL
! Default: non-BGP: 4, BGP: 1
! Max:  Probably 16
```

The max value can be checked with:

```
R# sh ip route summary
IP routing table name is Default-IP-Routing-Table(0)
IP routing table maximum-paths is 16
```

For static routes you can always add new routes up to the maximum number of routes.

### Load sharing with EIGRP

For EIGRP you can also use the **variance** command to do [Unequal Cost Load Balancing](https://nyquist.eu/eigrp-101/#7_Load_Balancing). EIGRP is the only IGP that can do Unequal Cost Load Balancing because the DUAL algorithm enables it to find loop-free paths to the same destination. That’s why only routes from Feasible Successors can be used for load balancing. Also, by default, EIGRP sends traffic on each path inversely proportional to the metric. To change this, use:

```
R(config-router)# traffic-share min across interface
! Default on RIP and OSPF
```

This command will make EIGRP loadbalance only on paths that have the minimum cost. The other paths will still be in the routing table, but will not receive traffic. To revert to the default, use:

```
R(config-router)# traffic-share balanced
! Default for EIGRP. Not available for OSPF and RIP
```

### Load sharing with BGP

BGP doesn’t support load balancing by default, but it can be enabled with the same command:

```
R(config-router)# maximum-paths [ibgp] VAL
```

It will do equal cost load balancing, but it can do it proportional to the bandwidth, with the following command:

```
R(config-router-af)# bgp dmzlink-bw
```

To actually work, the bandwidth value must be set and sent between neighbors using extended communities:

```
R(config-router-af)# neighbor NEIGH-ADDR activate
R(config-router-af)# neighbor NEIGH-ADDR send-community extended
R(config-router-af)# neighbor NEIGH-ADDR dmzlink-bw
```

### Load sharing in real life

In reality, the actual paths used can differ, based on the packet switching mode that is in use. Check [this article](https://nyquist.eu/how-cef-works/#Load_sharing) for more details


# How CEF works

## Process Switching

### How it works

1. Network interface detects a new packet on the wire. The interface will receive the packet and will place it in the I/O memory. It will then send a “receive interrupt” to the processor to indicate that a new packet needs to be switched. The interrupt itself contains the header information required to switch the packet.
2. Upon getting the “receive interrupt”, the processor now runs the appropriate process to deal with this packet (based on the header information). For IPv4 it’s *ip\_input*. The *ip\_input* process will lookup the destination in the routing table end will find the next hop address. With this address, it will perform a new lookup in the ARP cache and will find the information to rebuild the L2 header of the packet. Next, the *ip\_input* process will rewrite the packet’s L2 header with the new information, while it’s stored in the I/O memory. The packet is then queued for delivery.
3. The processor of the outgoing queue will take the packet from the I/O memory and will send it out the wire.
4. Once transmission is done, the interface processor will let the main processor know that the packet has been sent and the memory can be freed.
5. Upon receiving the information from the outgoing interface, the processor increments its packet coutners and frees the I/O memory where the packet used to be stored.

### Load sharing

Process switching performs per-packet load sharing. This means that if there are multiple paths to reach a destination, packets will be distributed evenly between the possible paths. For unequal cost routes (EIGRP), the distribution will be inversely proportional with the metric of each path. which means more packets will be sent over a path with a lower metric.

## Fast Switching

### How it works

Fast switching has 2 “modes on operation” depending if it switched a similar packet before (to the same destination) or not. When a packet for a destination is first received, the router will use process switching with the addition that once outgoing interface and next hop are determined, they are saved in a “route cache”. So first packet is always process switched, the rest are fast switched

1. Network interface detects a new packet on the wire. The interface will receive the packet and will place it in the I/O memory. It will then send a “receive interrupt” to the processor to indicate that a new packet needs to be switched. The interrupt itself contains the header information required to switch the packet. The interrupt will also look into the route cache to determine if the information needed to switch the packet already exists or not in the route cache. If there is no entry, the packet is process switched.
   * While performing the standard process switching, the *ip\_input* process will also update the route cache with the information needed to rebuild packets for the determined next hop. Subsequent packets with similar destination will be switched according to the route cache entry.
2. If a useful entry is found in the route cache, the interrupt software will rewrite the header and will also determine the outgoing interface.
3. The processor of the outgoing queue will take the packet from the I/O memory and will send it out the wire.
4. Once transmission is done, the interface processor will let know the main processor that the packet has been sent and the memory can be freed.
5. Upon receiving the information from the outgoing interface, the processor increments its packet counters and frees the I/O memory where the packet used to be stored.

Since the cache entries are added only for the first packet, there is a challenge to keep the cache updated with network changes. There are mechanisms that will invalidate entries based on network events (E.g when an ARP entry for a next hop changes, or when the route for a destination changes) and there is the cache aging process that will randomly invalidate a part of the cache (1/20th or 1/5th every minute).&#x20;

It’s understandable now, that in environments where the paths in the network are not stable enough, the caching mechanism will not be able to cope with the network changes, resulting in a lot of packets being process switched.

Intial versions of fast switching used hashed tables to store information, but later versions use a [radix tree](https://en.wikipedia.org/wiki/Radix_tree). The issue with using the radix tree, is that the subnet information is not used, so routes that share the same initial bits may get overlapped entries in the radix tree. To avoid such situations, IOS imposes some rules on how the tree is built. Further developments (Optimized switching and CEF) resolved this issue by using an [mtree](https://en.wikipedia.org/wiki/M-tree) instead of the radix tree.

### Load sharing

Using fast switching, even if there are multiple paths to a destination, a router will keep using the path determined by the first processed switched packet until that path is invalidated. This can happen even randomly, so there is a lack of determination in regards to how packets will be delivered. This may generate link polarization depending on how things work out.

### **Optimized switching**

Optimized switching is an improved version of fast switching that applies to IPv4 only. The improvement is mostly related to the way the route cache is organized, as a [mtree](https://en.wikipedia.org/wiki/M-tree) with a depth of 4, where each node has 256 children. This maps to the 4 bytes in an IP address, making it easy to store information for any destination. The other rules of Fast switching still apply.

## CEF – Cisco Express Forwarding

CEF works by building 2 structures:

* **CEF table,** aka **FIB** (Forwarding Information Base) – implements the routing table in an easy-to-search [mtree](https://en.wikipedia.org/wiki/M-tree) data structure. The leaves of the tree store pointers into the adjacency table. It is stored in TCAM tables. When a topology change occours, the routing table is updated and the changes are reflected in the FIB as well.
* **Adjacency table** – holds information regarding the L2 fields that needs to be overwritten over the incoming packet’s header and it is populated automatically from the ARP table or other L2 protcols.

Since entries in the CEF table and the adjacency table are all updated when the routing tabel or the L3 to L2 information tables (ARP, Framer Relay maps) change, the CEF information is always up to date and no aging is needed (as with Fast switching).

### CEF table (FIB)

The main difference between a route table and FIB is that when FIB entries are updated, their next-hop doesn't have any additional dependencies and already resolves to a next hop that is attached. \
If route for prefix A has the next-hop B and route to B has the next-hop C (direclty connected) then the FIB entry for A will show the next-hop C.

```
R# sh ip cef
Prefix               Next Hop             Interface
0.0.0.0/0            no route
0.0.0.0/8            drop
0.0.0.0/32           receive              
10.10.10.0/30        attached             Ethernet0/0
10.10.10.0/32        receive              Ethernet0/0
10.10.10.1/32        receive              Ethernet0/0
10.10.10.2/32        attached             Ethernet0/0
10.10.10.3/32        receive              Ethernet0/0
127.0.0.0/8          drop
192.168.100.0/24     attached             Ethernet0/1
192.168.100.0/32     receive              Ethernet0/1
192.168.100.1/32     receive              Ethernet0/1
192.168.100.255/32   receive              Ethernet0/1
192.168.110.0/24     10.10.10.2           Ethernet0/0
224.0.0.0/4          drop
224.0.0.0/24         receive              
240.0.0.0/4          drop
255.255.255.255/32   receive              

```

* no route - there's no route to this destination. Should only appear for default route
* drop - traffic for these destinations should be dropped. This may show up for "special" prefixes like 0.0.0.0/8, 127.0.0.0/8, 224.0.0.0/4, 240.0.0/4
* receive - traffic for these destinations is considered local by the host and will not be forwarded, but it should be further processed by the router. For example traffic destined to local IP addresses and to their respective broadcast addresses
* attached - traffic for these destinations is directly connected so traffic should be forwarded
* IP address - traffic for these destinations should be forwareded to the indicated next-hop IP address

CEF can be disabled per interface or globaly

```
R# show ip interface INTF | i CEF
  IP CEF switching is enabled
R(config)# [no] ip cef
R(config-if)# [no] ip route-cache c
```

When looking at the CEF table you can filter the output by outgoing interface or destination prefixes (with longer-prefixes match). Another handy tool when troubleshooting is the cef exact-route command that allows you to see how a packet will be forwarded:

```
R# show ip cef INTF
R# show ip cef PREFIX MASK [longer-prefixes]
R# show ip cef exact-route SRC-IP DST-IP
```

### Adjacency table

```
R# show adjacency
```

Entries in the adjacency table usually have information about outgoing interfaces and information to rebuild the packet, but there are some exceptions, that will make the packets to be process switched:

* Punt: move to the next switching layer for processing – special packets, or unsupported features
  * packets that have IP Header options
  * packets with an expiring TTL
  * packets forwarded to a tunnel interface
  * packets with unsupported encapsulation types or that are routed to interfaces with unsupported encapsulation types
  * packets that exceed the MTU value of the outgoing interface and that need to be fragmented
* Drop: drop the packet, but the prefix is checked
* Discard: drop the packet
* Null: packets for Null0 interface. Drop them
* Glean: an ARP request needs to be sent in order to “glean” L2 information

{% hint style="info" %}
<https://www.cisco.com/c/en/us/support/docs/ip/express-forwarding-cef/17812-cef-incomp.html#types>
{% endhint %}

### Load sharing

Since the data structure for the CEF table doesn’t hold actual L2 information, but rather pointers to the adjacency table, CEF is able to point some destinations to single entries in the adjacency table, while pointing other entries to a “load-share table”. When reaching the load-share table, one entry will be selected based on the load-sharing algorithm that is used (per destination or per packet). This entry will then point to one entry in the adjacency table, which will be used to rebuild the packet.\
For per destination load balancing a hash is computed out of the source and destination IP address. This hash points to exactly one of the adjacency entries in the adjacency table, providing that the same path is used for all packets with this source/destination address pair. If per packet load balancing is used the packets are distributed round robin over the available paths, similar to how process switching worked. The default mode is per destination, but this can be changed using:

```
R(config-if)# ip load-sharing {per-packet|per-destination}
!Default: per-destination 
```

When you see traceroute packets being load-shared per-packet, it is because ICMP packets are process-switched, so they do not fall into the CEF-switched traffic.

### Centrialized vs Distributed CEF

For devices that run Centralized CEF, the FIB and the Adjacency table reside on the Route Processor (RP) and the RP performs CEF forwarding.

With Distributed CEF, the RP maintains the FIB and the Adjacency table but there is also a process(IPC - InterProcess Communication) that synchronizes the RP FIB and Adjacency table to with the line cards. This way the CEF forwarding is performed by the line cards directly.


# Routing Order of Operations

The original information was taken from Cisco article on [NAT Order of Operations](https://www.cisco.com/en/US/tech/tk648/tk361/technologies_tech_note09186a0080133ddd.shtml). However, this order helps understand other features, like WCCP.

1. If IPSec then check input access list
2. decryption – for CET (Cisco Encryption Technology) or IPSec
3. check input access list
4. check URPF (Unicast Reverse Path Forwarding)
5. check input rate limits
6. input accounting – update stats
7. redirect to web cache
8. If configured with ip nat outside: NAT outside to inside (global to local translation)
9. policy routing
10. routing
11. If configured with ip nat outside: NAT inside to outside (local to global translation)
12. crypto (check map and mark for encryption)
13. check output access list
14. inspect (Context-based Access Control (CBAC))
15. TCP intercept
16. encryption
17. Queueing


# NSF – Non Stop Forwarding

## What is NSF

NSF is a feature that allows routers to keep on forwarding traffic (non stop forwarding) even in the event of a restart. This is done by separating the control and the data plane, having one process involved in building the routing table and another process in forwarding the packets.This feature takes advantage of CEF which updates the line cards with the information from FIB.

In order for NSF to work, routers must be NSF-capable or NSF-aware. A NSF-capable router is a router that can perform restarts without disrupting packet forwarding, while a NSF-aware router understands NSF-specific signaling from NSF-capable routers.\
NSF requires these additional features because, while performing restart, a router will not be able to send Hellos to its peers. This normally should result in the neighbor relationship being dropped, routes being lost, packets being dropped.\
Routers exchange NSF capabilities information with each other and when a NSF-capable router performs a restart, its NSF-aware peers will change the default behavior in order to prevent breaking the neighbor relationship.\
NSF-aware routers are also called NSF-helpers because they will help a router performing a NSF restart to re-sync with the network as soon as possible.

## NSF Support

### EIGRP

To configure NSF supports for EIGRP, use:

```
R(config)# router eigrp AS-NUMBER
R(config-router)# address-family ipv4 unicast [vrf VRF] AS-NUMBER
R(config-router-af)# nsf
! Disabled by default
```

To modify the EIGRP NSF timers, use:

```
R(config-router-af)# timers graceful-restart purge-time SEC
! Default: 240 sec
R(config-router-af)# timers nsf converge SEC
! Default: 120 sec
R(config-router-af)# timers nsf route-hold SEC
! Default: 240 sec
R(config-router-af)# timers nsf signal SEC
! Defualt: 20 sec
```

### OSPF

Cisco routers support 2 NSF modes, a Cisco-proprietary and an IETF-open-standard version (RFC 3623).

**2.2.1 Cisco Mode**

To configure NSF in Cisco Mode:

```
R(config)# router ospf PROC-ID
R(config-router)# nsf cisco enforce global
! Enables a router to be NSF-capable. Disabled by default
R(config-router)# nsf helper [disable]
! Enables a router to be NSF-aware. Enabled by default
```

**2.2.2 IETF Mode**

To configure NSF for IETF Mode:

```
R(config)# router ospf PROC-ID
R(config-router)# nsf ietf [restart-interval SEC]
! default: 120 sec
R(config-router)# nsf ietf helper [disable]
!NSF-aware mode is enabled by default.
```

Strict LSA Checking is a feature of the IETF NSF mode for OSPF. When a router acts as a NSF-aware router (NSF-helper) it will stop the nsf-helper process if it detects a change of a LSA that would be forwarded to the restarting router. To enable this feature, use:

```
R(config-router)# nsf ietf helper strict-lsa-checking
! disabled by default
```

**2.2.3 Verify OSPF NSF**

To verify how OSPF works with NSF, you can use:

```
R# show ip ospf neighbors
R# show ip ospf nsf
R# debug ospf nsf [detail]
```

#### 2.3 BGP

To enable NSF for BGP, use:

```
R(config)# router bgp AS-NUMBER
R(config-router)# bgp graceful-restart [restart-time SEC | stalepath-time SEC] 
! Default restart-time : 120 sec
! Default stalepath-time: 360 sec
```

This command is enabled per address-family. You can enable it for all address families with:

```
R(config-router)# bgp graceful-restart all
```

To verify:

```
R# show ip bgp neighbors [NEIGH-IP]
```


# RIP


# RIP 101

## Starting the routing process

```
R(config)# router rip
R(config-router)# network NETWORK-ADDR
```

In RIP there is no wildcard option when configuring the network command. The NETWORK-ADDR will always be considered a classful address and will match all interfaces that are in the same classful network:

```
R(config-router)#network 172.16.23.45
R1(config-router)#do sh run | se router
router rip
network 172.16.0.0
```

The network command enables RIP on the interfaces where an IP address is configured that is part of the classful network that is defined in the command. By enabling RIP on that interface, the router will advertise the subnets for the interfaces where RIP has been enabled.

### RIP Versions

* By default, a router will send V1 updates, but will accept both v1 and v2
* Version 2 can be enabled globally or per interface:

  ```
  ! Globally, inside the router process
  R(config-router)# version {1|2}
  ! Per interface
  R(config-if)# ip rip {send|receive} version {1|2|1 2}
  ```

#### **RIPv1 vs RIPv2**

* RIPv2 is backwards compatible with RIPv1, but RIPv1 ignore version 2 updates
* RIPv2 supports VLSM, RIPv1 doesn’t – RIPv1 is a classful routing protocol that does not send netmask information in routing updates
* RIPv2 sends updates using multicast 224.0.0.9 instead of broadasts like RIPv1
* RIPv2 supports authentication (MD5 and text)
* RIPv2 supports route tagging (for redistribution)

#### **RIPv1 updates**

RIPv1 does not send netmask information in routing updates, so it uses a system to deduce the mask associated with a destination.\
When sending an update:

* Is the advertised network part of the same major network as the outgoing interface?
  * **NO**: Advertise the major network of the route (Auto-Summary)
  * **YES**: Does the advertised network has the same netmask as the outgoing interface?
    * **YES**: Advertise the network
    * **NO**: Is it a host route? (/32 netmask)
      * **YES**: advertise the network
      * **NO**: Do not advertise the network

When receiving an update:

* Is the received network part of the same major network as the incoming interface?
  * **YES**: Does it has non-zero bits in the host part (Host network)?
    * **YES**: Add it to the routing table as a host network (/32)
    * **NO**: Add it to the routing table with the same netmask as the incoming interface
  * **NO**: Then it must be a major netwokr. Are there any subnets of this major network already in the routing table, received on other interfaces?
    * **YES**: Drop the route (Does not support discontiguous networks)
    * **NO**: Apply a classful mask to the network address and add it to the routing table

### Split Horizon

* “Split Horizon” rule blocks information about routes from being advertised out the interface where that information was received on.
* “Split Horizon” is on by default, except on Frame Relay physical interfaces. It should be manually disabled on multipoint subinterfaces, or even on Ethernet interfaces, when it is needed.
* To disable/enable split-horizon on an interface use:

  ```
  R(config)# [no] ip split-horizon
  ```
* To verify if split-horizon is enabled or not, use:

  ```
  R# show ip interface INTERFACE | i Split
   Split horizon is enabled
  ```

### Passive interfaces

Setting an interface as passive will disable sending advertisements, but will not disable receiving advertisements.

```
! Use this command to make all interfaces passive by default:
R(config-router)# passive-interface default
! Use this to enable/disable sending on one interface:
R(config-router)# [no] passive-interface INTERFACE
```

## Neighbors

### Static Neighbors

Static neighbors for RIP can be defined with the command:

```
R(config)# neighbor NEIGH-ADDR
```

Now, communication with this neighbor will be done using unicast addresses. Unlike EIGRP, multicast addresses will still be used on the interfaces where static neighbors are defined.

### Authentication

RIPv2 supports both MD5 and text authentication.

1. Define the key chain:

   ```
   R(config)# key chain KEY-CHAIN
   R(config-keychain)# key KEY-NUMBER
   R(config-keychain-key)# key-string KEY-STRING
   ! Optionally, define an accept-lifetime
   R(config-keychain-key)# accept-lifetime START-TIME {infinte|END-TIME|duration SEC}
   ! Optionally, define the a send-lifetime
   R(config-keychain-key)# send-lifetime START-TIME {infinte|END-TIME|duration SEC}
   R(config-keychain-key)# end
   ```
2. Apply it on an interface:

   ```
   ! set the authentication as MD5 or plain text
   R(config-if)# ip rip authentication mode {md5|text}
   R(config-if)# ip rip authentication key-chain KEY-CHAIN
   ```

When using plain text authentication, the Key ID is ignored. Whe using md5 authentication, the Key ID should match for a properly working network. But in fact, a router will accept RIP packets authenticated with a lower key ID and will reject RIP packets authenticated with a higher key ID.

## Timers

* RIP sends routing updates at regular intervals and when the network topology changes.
* Routing updates are sent every 30 sec(default UPDATE)
* If no update is received for a route for 180 sec(default INVALID), the route is marked inaccessible and advertised as unreachable (metric 16). However, the route is still used for forwarding packets
* If no update is received for a route for 240 sec(default FLUSH), the route is removed from the routing table.
* When a router receives an update with a metric worse than the one it already has, it puts the route in **hold-down** state. The route stays in this state for 180 sec (default HOLD-DOWN), interval during which outing information regarding better paths is suppressed. The route is marked inaccessible and advertised as unreachable. However, the route is still used for forwarding packets. When HOLD-DOWN expires, routes advertised by other sources are accepted and the route is no longer inaccessible.

The timer values can be changed globally:

```
! Globally:
R(config-router)# timers basic UPDATE INVALID HOLD-DOWN FLUSH [SLEEP]
! Default: 30 sec, 180 sec, 180 sec, 240 sec
```

The UPDATE interval can be changed per interface:

```
R(config-if)# ip rip advertise UPDATE
! Default: 30 sec
```

## Packets

* RIP uses UDP 520 (both source and destination) to exchange routing information
* One RIP update packet can accomodate up to 25 routing updates
* RIPv1 sends updates as broadcasts to 255.255.255.255
* By default, RIPv2 sends updates to the multicast address 224.0.0.9 (MAC: 01-00-5E-00-00-09)
* Multicasts in RIPv2 can be sent as broadcasts too, using the commnand:

  ```
  R(config-if)#ip rip v2-broadcast
  ```
* RIPv2 can also send updates as unicast if the neighbor is defined using:

  ```
  R(config)# neighbor NEIGHBOR-ADDR
  !Unless the interface is passive, the router will send both multicast
  !and unicast address to the statically defined neighbors
  ```
* To check what kind of updates are sent, use:

  ```
  R# debug ip rip
  ```
* Updates are sourced only from the primary IP of the interface.
* When an update is received, the router checks if the source IP is in the same subnet as any IP address on the receiving interface. If it doesn’t find an IP address on the interface in the same subnet as the update source, it drops the update.
* It can be a problem when using PPP or unnumbered links, because each end of the link can be in different subnets.
* You can disable this behavior using:

  ```
  R(config-router)# no validate-update-source
  ```
* Normally, updates are sent every UPDATE , but you can set RIP to only send updates when there is a change in the topology. This can be usefull on Dial interfaces

  ```
  R(config-if)# ip rip triggered
  ```

## RIP Metric

* RIP uses hop count as route metric
* A metric of 16 equals infinty – The route is considered inaccesible

### Offset lists

Use offset lists to add a value to the calculated metric:

```
R(config-router)# offset-list ACL {in|out} OFFSET
```

Offset value must be between 1 and 16. Remember that any route that has a metric of 16 is considered down. You can control how far a RIP update can reach by advertising the routes with a certain metric.\
E.g., advertising the routes with a metric of 15, will allow only the next hop router to accept the route. Provided no other metric manipulation is done, other routers, beyond the next hop router, will consider the route inaccessible.

## RIP Administrative Distace

The default AD for RIP is 120. This can be changed using the distance command:

```
R(config-router)# distance AD [SRC-ADDR WILDCARD [ACL]]
```

## Load Balancing

RIP can only perform load balancing when the metric is equal. Use offset lists to modify the route metric in order to manually engineer load balancing.

## Distribute lists

Use them to filter the updates that are sent or received

```
! Using ACLs
R(config-router)# distribute-list ACL {in|out} [INTERFACE]
! Using Prefix-list
R(config-router)# distribute-list prefix PREFIX-LIST {in|out} [INTERFACE]
! Using gateway - filter based on the source of the update:
R(config-router)# distribute-list gateway PREFIX-LIST1 [prefix PREFIX-LIST2] {in|out} [INTERFACE]
```

## Summarization

By default, RIP summarizes at major network boundaries. This can be disabled using:

```
R(config-router)# no auto-summary
```

Manual summarization can be done per interface, using:

```
R(config-if)# ip summary-address rip SUMMARY-ADDR SUMMARY-MASK
```

When manually summarizing, you can’t go past the major network. As a workaround, you can add a static route to NULL and redistribute static. The summary route will have a metric equal to the lowest metric of all its children. Unlike EIGRP, RIP does not add a NULL route when summarizing.

## Default routes

You can enable RIP to advertise a default route using:

```
R(config-router)# default-information originate [route-map ROUTE-MAP]
```

The route map can be used to make conditional advertisement of the default route:

* Advertise a default route only on some interfaces:

  ```
  R(config)# route-map ROUTE-MAP
  R(config-route-map)# set interface INTERFACE
  ```
* Advertise a default route only if another route is in the routing table:

  ```
  R(config)# route-map ROUTE-MAP
  R(config-route-map)# match ip address prefix-list PREFIX-LIST
  ```


# EIGRP


# EIGRP 101

## Starting the routing process

```
R(config)# router eigrp AS-NUMBER
! AS Number mast match between neighbors
R(config-router)# network NETWORK-ADDR [WILDCARD]
!If no wildcard is specified, the network is considered classful
```

* EIGRP will advertise routes learned by the EIGRP process and all routes that appear directly connected on the interfaces that are matched by the network command. These include static routes that point to an interface and that are matched by a network command. These routes are considered directly connected and are redistributed as internal routes.
* secondary IP addresses are not advertised when using the network command. They can only be redistributed into EIGRP
* With Split Horizon enabled(default), it will not advertise a route back on the outgoing interface of that route.
* Routes that don’t make it into the routing table are not advertised

To see the interfaces that run EIGRP use:

```
R3#sh ip eigrp interfaces
IP-EIGRP interfaces for process 345

                        Xmit Queue   Mean   Pacing Time   Multicast    Pending
Interface        Peers  Un/Reliable  SRTT   Un/Reliable   Flow Timer   Routes
Fa0/0              1        0/0        71       0/2          384           0
Se1/0.301          0        0/0         0       0/1            0           0
Lo0                0        0/0         0       0/1            0           0
```

### Split Horizon

When Split Horizon is enabled on an interface, Update and Query packets are not sent for destinations which have this interface as outgoing. This could be a problem in Hub and Spoke Frame Relay topologies. Use this command to disable Split Horizon:

```
R(config-if)# no ip split-horizon eigrp AS-NUMBER
```

Make sure you know the difference between the RIP and the EIGRP command that disables Split Horizon:

```
! EIGRP:
R(config-if)# no ip split-horizon eigrp AS-NUMBER
! RIP:
R(config-if)# no ip split-horizon
```

Some IOS implementations disable split-horizon on interface where “encapsulation frame-relay” is configured.

### Passive interfaces

When defining a passive interface, the EIGRP process will advertise the network but will not send or accept EIGRP messages on the interface.

```
R(config-router)# passive-interface INTERFACE
```

This behavior is different than RIP’s, where a passive interface would still accept RIP advertisements.

You can also enable passive interfaces by default and then disable it on the interfaces where EIGRP should run:

```
R(config-router)# passive-interface default
R(config-router)# no passive-interface INTERFACE
```

## Neighbors

When a router running EIGRP receives valid HELLOs from another router, it adds it to the neighbor list. To see the neighbors list use:

```
R# show ip eigrp neighbors [detail]
```

A HELLO message contains the AS number and the K values of the router sending the message. In order to become neighbors, two routers must share the following values:

* K values
* AS number
* Primary subnet
* Authentication

By default, when valid HELLOs are not received for an entire HOLD-TIME period, the router considers the neighbor to be down. See [Timers ](#timers)section below.

### Static Neighbors

When a neighbor is defined, all communication with it is done using unicast packets:

```
R(config-router)# neighbor NEIGH-ADDR OUT-INTERFACE
```

Unlike RIP, this will disable multicast EIGRP on the interface, so no dynamic neighbors will be discovered.\
This config should be used on Frame Relay Hub & Spoke networks when two spokes should become neighbors.

### Authentication

EIGRP supports only MD5 authentication of EIGRP messages:

1. Define the key chain

   ```
   (config)# key chain KEY-CHAIN
   R(config-keychain)# key KEY-NUMBER
   R(config-keychain-key)# key-string KEY-NAME
   ! Optionally, define the an accept-lifetime
   R(config-keychain-key)# accept-lifetime start-time {infinte|END-TIME|duration SEC}
   ! Optionally, define the an send-lifetime
   R(config-keychain-key)# send-lifetime start-time {infinte|END-TIME|duration SEC}
   ```
2. Apply it on the interface

   ```
   R(config-if)# ip authentication mode eigrp AS-NUMBER md5
   ! sets the authentication to MD5
   R(config-if)# ip authentication key-chain eigrp AS-NUMBER KEY-CHAIN
   ```

When sending EIGRP messages, the router uses the lowest key number among all current valid keys. When receiving EIGRP messages, the router checks the MD5 digest using all current valid keys. Both the key ID and the Key-string must match in order to form an adjacency.

## Timers

### Hello interval

It specifies how often a router sends EIGRP HELLO packates. The default timer is:

* **5 seconds** for almost all interfaces
* **60 seconds** for Frame Relay physical interfaces or multipoint subinterfaces with a bandwidth lower than T1(1544kbps)

The default value can be changed using:

```
R(config-if)# ip hello-interval eigrp AS-NUMBER SEC
```

To verify, use:

```
R# show ip eigrp interface INTERFACE detail
IP-EIGRP interfaces for process 145

                        Xmit Queue   Mean   Pacing Time   Multicast    Pending
Interface        Peers  Un/Reliable  SRTT   Un/Reliable   Flow Timer   Routes
Fa0/0              1        0/0        48       0/2          240           0
  Hello interval is 5 sec
  Next xmit serial <none>
  Un/reliable mcasts: 0/1  Un/reliable ucasts: 2/4
  Mcast exceptions: 1  CR packets: 1  ACKs suppressed: 0
  Retransmissions sent: 1  Out-of-sequence rcvd: 0
  Authentication mode is not set
  Use multicast
```

### Hold Time

If a router does not receive Hello Messages for an entire Hold Time, the router considers the neighbor to have failed. The default timer is 3xHELLO-INTERVAL:

* **15 seconds** for almost all interfaces
* **180 seconds** for Frame Relay physical interfaces or multipoint subinterfaces with a bandwidth lower than T1(1544kbps)

The default value can be changed using:

```
R(config-if)# ip hold-time eigrp AS-NUMBER SEC
```

### Active Timer

When the router sends a Query, it will wait the time specified in the active timer for a replies from its neighbors. If no reply is received in the specified interval, the route is declared dead:

```
R(config-router)# timers active-time {ACTIVE-TIME|disabled}
! Default: 180 sec
```

See [Going Active](/ipv4/ipv4-routing/eigrp/more-eigrp-features#going-active-sending-queries) for details.

## Packets

EGIRP uses IP protocol 88 (RTP=Reliable Transport Protocol). EIGRP uses both unicast and mulitcast packets. Except for HELLO and ACK packets, the other packets require ACK from the neighbors. A router would retry 16 times to send a packet before neighbor relationship is reset. All packets are sourced from the primary IP address of the interface.

* **HELLO** packets are sent every HELLO-INTERVAL as multicasts to 224.0.0.10 or as unicasts to each static neighbor. They contain the K values used by the router as well as the HOLD-TIME – how much time to wait for a HELLO, before resetting adjacency. EIGRP packets are sourced from the primary address of each interface.
* **UPDATE** packets are sent as unicast when neighbors are discovered initially and as multicasts whne updates are genertated by network changes. For each route, a packet contains the prefix, prefix length, hop count, and components of the advertised metric (Bandwidth, Load, Delay, Reliability, MTU). These packets are sent only when changes occur and only to the routers that need the update. EIGRP doesn’t use periodic updates. Update packets need to be acknowledged.
* **QUERY** packets are used when “going active” – see [Going Active](/ipv4/ipv4-routing/eigrp/more-eigrp-features#going-active-sending-queries). QUERY packets are sent as multicast and need to be ACKed.
* **REPLY** packets are used to reply to QUERY packets. They are sent as unicasts and need to be ACKed.
* **ACK** packets are sent as an acknowledgement to Query and Reply messages. ACK are always unicasts, and EIGRP expectes one ACK from each neighbor.
* **GOODBYE** packets are sent when the EIGRP process is shut down or restarted to inform the neighbors.

If there are many changes in the network, EIGRP messages can overwhelm a link. Use this command to limit the bandwidth used by EIGRP Updates per interface:

```
R(config-if)#ip bandwidth-percent eigrp AS-NUMBER BW-PERCENT
!default: BW-PERCENT = 50
```

## EIGRP metric

See [here](#eigrp-metric) for information about EIGRP metric and offset lists.

## EIGRP Administrative Distance

By default, EIGRP AD is **90** for internal routes, **170** for external routes and **5** for summary routes. The default can be changed using:

```
R(config-router)# distance eigrp INTERNAL-AD EXTERNAL-AD
```

The AD can be changed per routing source and destination using:

```
R(config-router)# distance AD DESTINATION-IP WILDCARD-MASK [ACL]
```

## Load Balancing

In addition to Equal Cost Load Balancing, EIGRP can perform Unequal cost load balancing. To enable this, the variance command must be used.\
When setting variance, all routes that have a metric lower then the **FD \* VARIANCE** are added to the routing table and traffic can be load balanced between them.

```
R(config-router)# variance VAR
```

## Distribute Lists

Used to filter updates going out or coming into the EIGRP process

```
! Using ACLs
R(config-router)# distribute-list ACL {in|out} [INTERFACE]
! Using Prefix-list
R(config-rotuer)# distribute-list prefix PREFIX-LIST {in|out} [INTERFACE]
! Using gateway - filter based on the source of the update:
R(config-router)# distribute-list gateway PREFIX-LIST1 [prefix PREFIX-LIST2] {in|out} [INTERFACE]
! Using route maps:
R(config-router)# distribute-list route-map ROUTE-MAP {in|out} [INTERFACE]
```

## Summarization

### Auto Summarization

By default, EIGRP performs an auto-summarization each time it crosses a border between two different major networks. To disable this behaviour and to advertise the component routes, use:

```
R(config-router)# no auto-summary
```

The auto-summary routes appear as internal routes and the metric of the summary route is the best metric from among the summarized routes. On the router doing the summarization, a route to Null0 is added for the summarized address, so that traffic that is destined for other destinations than the component routes, but inside the same major network, is discared.

EIGRP will not auto-summarize external routes unless there is a component of the same major network that is an internal route

### Manual Summarization

Manual sumamrization can be done on an interface, with no limitation of the bit boundary, using:

```
R(config-if)# ip summary-address eigrp AS-NUMBER NETWORK-ADDR NETWORK-MASK [AD] [leak-map LEAK-MAP]
```

The router will advertise the summary address instead of any component, as long as there is one component in the routing table.

By default, manual summary routes have an AD of 5 and point to Null0. Due to the low AD value, they can override other learned routes with the same prefix (like a default route). Use the concept of a floating summary route to change the default AD when defining the manual summary. Use a value of 255 to stop the summary route from getting into the routing table

Leak maps enable sending routing information for specific routes defined by a route-map, and the summary route for all other.

## Default routes

### Redistribute Static Route to 0.0.0.0 into EIGRP

```
R(config)# ip route 0.0.0.0 0.0.0.0 NEXT-HOP
R(config)# router eigrp AS-NUMBER
R(config-router)# redistribute static
```

### Define a default network

By default 0.0.0.0/0 is considered the default route, but EIGRP can advertise another route as the default route as long as it is marked as the default-network

```
R(config)# ip default-network NETWORK-ADDR
! NETWORK-ADDR is classless
```

If we have the NETWORK-ADDR with its default class (A/B/C) in our routing table, then the next step is to advertise it into EIGRP. This can be done via redistribution (of connected/static routes) or the network command (provided the matched interfaces will bring the specific prefix into the EIGRP topology).\
Things get ugly when the default-network is not int our routing table with the default class. Let’s take an example:\
If we have a loopack address with the network 5.5.5.5/24 and we want it to become the default network, then we will have to run the following commands:

```
R(config)# ip default-network 5.0.0.0
R(config)# ip default-network 5.5.5.5
```

The second command will generate a static route in the config, that will be deleted only when the default-network is removed from the config:

```
R(config)#sh run | i ip route
ip route 5.0.0.0 255.0.0.0 5.5.5.5
```

Now we have a route pointing to the Class A network 5.0.0.0 via the next-hop 5.5.5.5. We can add the 5.0.0.0 network to EIGRP via redistribution (static) to EIGRP or add the 5.5.5.5 interface to the network list. From now on, EIGRP will advertise a route to 5.0.0.0 that is also marked as a candidate default.

### Manual Summary to 0.0.0.0/0

```
R(config-if)# ip summary-address eigrp 0.0.0.0 0.0.0.0 [AD]
```

Use a lower AD not to override other default routes learned by the downstream routers.


# EIGRP Metric

$$
Metric\_{EIGRP} =\begin{cases}\[(K\_1B+\frac{K\_2B}{256-L}+K3D)\*{\frac{K\_5}{R+K\_4}}]\*256 & K\_5!=0\ (K\_1B+\frac{K\_2B}{256-L}+K3D)\*256 & K\_5=0\B+D & Default \end{cases}
$$

\
you will have to round down to the nearest integer

* default K values: **(K1,K2,K3,K4,K5) = (1,0,1,0,0)**
  * default EIGRP metric: **B+D**
  * remember it as **BuiLDeR 20**
* The EIGRP metric is 256 times larger than IGRP metric.

## Metric Components

* **Bandwidth – B**
  * inverse lowest bandwidth along the path in kbps, scaled by 10^7. The router receives the bandwidth "along the path" in advertisements from other routers. Therefore loweste bandwidth along the path will be the minimum value between the link's bandwidth and the advertised bandwidth.
  * Interfaces with bandwidth higher than 10Gbps (10^7 kbps) are considered similar from EIGRP’s metric standpoint (B=1). See Wide-Metrics if you need Bandwidth to account for interfaces over 10Gbps.

$$
B = \frac{10^7}{\min\_{path}(Bw\[kbps])} = \frac{10^7}{\min(Bw\_{Link}, Bw\_{Advertised})}
$$

* **Load – L**
  * Highest load along the path
  * Dynamically determined by the router. It has values starting from 1 (no usage) to 255(fully utilized)

$$
L = \max\_{path}(Load) = \max(Load\_{Link}, Load\_{Advertised})
$$

* **Delay – D**
  * cumulative delay along the path in 10s of microseconds
  * EIGRP uses Delay to signal an unreachable route, by using the Delay value of 0xFFFFFF

$$
D = \sum\_{path}(Delay\[10 \mu s]) = Delay\_{Link}+Delay\_{Advertised}
$$

* **Reliability – R**
  * Lowest reliability along the path
  * Dynamically determined by the router as the percentage of successfully received packets on the interface scaled by 255. The maximum value of 255 means 100% reliability.

$$
R=\min\_{path} (Reliability) = \min(Relability\_{Link}, Reliability\_{Advertised})
$$

* **MTU**
  * MTU is used as a tiebreaker if the metric is the same for more paths
  * Largest MTU wins the tiebreak

$$
MTU = \min\_{path}(MTU\_{interface})=\min(MTU\_{Link}, MTU\_{Advertised})
$$

* **Hop Count**
  * Not used in the actual metric, but the value is passed from router to router
  * There is a default limit of 100 to the number of hops to a destination, but this can be changed up to 255.

$$
Hop\_{count}=\sum\_{path}(Hops) = Hops\_{Advertised} + 1
$$

When route information for a destination is received, it also contains the metric parameters used by the advertising router: Bandwidth, Delay, Load, Reliability, MTU. The receiving router compares the received information with the data it has for the incoming interface so it can find the lowest Bandwidth along the path, the highest Load along the path, the sum of Delays along the path, the lowest Reliability along the path and the lowest MTU along the path. Then, it can apply the formula to find the metric value for each path. The router that advertised the best path is considered the **Successor**, and the metric for that path is known as **„Feasible Distance” (FD)**

For each destination, the router also applies the formula on the parameters it received from the advertising router to calculate the metric from that router towards the destination. This is known as the **Advertised Distance (AD)** or **Reported Distance (RD).**

Using the DUAL algortithm, EIGRP will consider all paths that have a RD\<FD as loop free backup paths and will call the routers advertising them Feasable Successors(FS) for the route. The FS will be kept in the topology table and when the route through the successor fails, the router will imediately use the best FS. For all other routes, DUAL will not consider them as backup paths, because they can be part of a routing loop (even if they aren’t actually). The idea is that if RD>FD, then that router is closer to the destination then us and it should route through us. If we use it as our next hop then it could send it back to us resulting in a routing loop.

## How to modify metric

### Changing metric components

* **Bandwidth**

  ```
  R(config-if)# bandwidth BANDWIDTH
  ! in kbps
  ```
* **Delay**

  ```
  R(config-if)# delay DELAY
  ! in 10s of usec
  ```
* **Load** and **Reliability** cannot be changed manually. They are caculated by the router based on a 5 minute average. The values can be seen using:

  ```
  R# show interface INTERFACE
  R3#sh int fa0/0 | i MTU|load
    MTU 1500 bytes, BW 10000 Kbit, DLY 1000 usec,
       reliability 255/255, txload 1/255, rxload 1/255
  ```

### Changing K values

```
R(config-router)# metric weights TOS K1 K2 K3 K4 K5
```

Default values for (K1,K2,K3,K4,K5) is (1,0,1,0,0) and only 0 is supported for TOS. K values must match between neighbors.

### Using offset lists

```
R(config-router)# offset-list ACL {in|out} OFFSET
```

Adds the OFFSET to the metric. It can be used to create multiple equal cost links.

## Maximum Hops

Even though EIGRP doesn’t use hop count in it’s metric calculation, it still counts how many hops away is the router that first advertised this network. You can see the hop count value using:

```
R1#show ip eigrp topology 4.4.4.0/24
IP-EIGRP (AS 10): Topology entry for 4.4.4.0/24
State is Passive, Query origin flag is 1, 1 Successor(s), FD is 2323456
Routing Descriptor Blocks:
24.0.0.2 (Serial1/0), from 24.0.0.2, Send flag is 0x0
Composite metric is (2323456/409600), Route is Internal
Vector metric:
Minimum bandwidth is 1544 Kbit
Total delay is 26000 microseconds
Reliability is 255/255
Load is 1/255
Minimum MTU is 1500
Hop count is 2
```

You can invalidate routes that are more than a number of hops away by running:

```
R(config-router)# metric maximum-hops HOPS
!Default: 100
```

All routes that have a hop-count greater than the maximum-hops, will not be added to the routing table.

## Wide metric

Newer IOS implementations support wide metrics, which can differentiate between interfaces higher than 10Gbps. [EIGRP Wide Metric](https://www.cisco.com/c/en/us/td/docs/ios-xml/ios/iproute_eigrp/configuration/xe-3s/ire-xe-3s-book/ire-wid-met.html) uses the following formula:

$$
WideMetric\_{EIGRP} = (K\_1B+\frac{K\_2B}{256-L}+K\_3D+K\_6E)\*\frac{K\_5}{R+K\_4}
$$

$$
Metric\_{EIGRP} =\begin{cases}(K\_1T+\frac{K\_2T}{256-L}+K\_3D+K\_6E)\*{\frac{K\_5}{R+K\_4}}& K\_5!=0\ K\_1T+\frac{K\_2T}{256-L}+K\_3D+K\_6E& K\_5=0\T+D & Default\end{cases}
$$

where&#x20;

* T = Throughput = $$\frac{10^7}{Bandwidth}\*65536$$
* D = Latency = $$\begin{cases}\frac{Delay}{10}*65536 & if Bw < 1Gbps\\\frac{10^7*\frac{65536}{10}}{Bw} & if Bw >= 1Gbps\end{cases}$$
* E = Extended Attributes.&#x20;

See  for details. Supporting Wide metric, requires changes in the EIGRP packets to add the additional information. For this reason, a router will send packets for both standard metric and wide metric, so additional bandwidth will be used. If it detects that all neighbors on an interface support wide metrics, then it will only send this version.

When the wide metrics are used, the metric can become larger than the maximum value allowed by the RIB so it needs to be scaled down using the command&#x20;

```
R(config-eigrp)# metric rib-scale
```

When it is cofigured, all EIGRP routes in the RIB are cleared and replaced with the new metric values.

<br>


# More EIGRP Features

## Router ID

The router ID is a 32 bit number, usually represented as a dotted decimal (like an IP address). The Router ID is determined when the routing process is started by the following algorithm:

1. Use the configured value

   ```
   R(config-router)# eigrp router-id ROUTER-ID
   ```
2. Use the highest up/up loopback ip address
3. Use the highest up/up non-loopback ip address

The router ID should be different between neighbors. When a router receives a routing update from a router with the same Router ID, it will ignore the update considering it is from itself.

## DUAL

The following information can be found looking at the EIGRP topology table:

```
R# show ip eigrp topology
```

### Feasible Distance and the Successor

A router uses the EIGRP formula to calculate the metric for each destination for all available paths.\
For each destination, an EIGRP router calculates the metric based on information it has from the neighbors that advertise the destination and its incoming interfaces. See [EIGRP Metric](/ipv4/ipv4-routing/eigrp/eigrp-metric). The best metric for each destination is called the **Feasible Distance**. This indicates the best path to reach the destination, and the router that is the next hop in this path is called the **Successor**.

### Reported Distance and Feasible Successors

For each destination, an EIGRP router also calculates a Reported Distance (RD), aka Advertised Distance, for the neighbor advertising it. It is actually the metric to the destination seen from the point of view of the neighbor. This will be lower than the metric calculated by the router.\
All routers that advertise a RD lower than the FD for each destination, are called **Feasible Successors**. These routers, are guaranteed to have a loop free path to the destination. EIGRP achieves quick conversion by immediately using one of the FS routes as the new Successor, in case the path via the Successor is down. All other paths cannot be guaranteed to be loop free and are not used in this step.

### Going Active – Sending Queries

What happens if there are no available FS in the topology table? The route will be marked as Active, and the router will send Queries to all its neighbors (except the Successor that just failed), asking if they have routes for that destination.

The neighbors will have to respond to the Query with a Reply, like this:

| Query from                       | Route state               | Action                                                                                                                                              |
| -------------------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| Successor                        | Passive                   | Attempt to find a new successor in the FS list. If none found, mark the destination active, and query all neighbors, except the previous successor. |
| Neighbor (not current Successor) | Passive                   | Reply with current Successor information                                                                                                            |
| Neighbor (not current Successor) | not in the topology table | Reply that the destination is unreachable                                                                                                           |

Waiting for each neighbor to reply, can put the route in a SIA (Stuck in Active) state. When a neghbor doesn’t reply in a resonable interval, a router will bring down the relationship with the neighbor. The default interval is 180 seconds and can be changed with:

```
R(config-router)# timers active-time {SEC|disabled}
```

To prevent unnecesary clearing of neighbor relationships, EIGRP sends at half the active-timer (90 sec by default) a SIA-Query message. If the router is still waiting for replies from its neighbors, it will send a SIA-Reply, so the active timer is reset. If the router does not receive a SIA-Reply, thent the neighbor relationship will be broken whent the active-timer expires.\
The Query domain can be limited by using route summarization and Stub Routers. The route will reply with an unreachable message if it has a route for a summary address, but not for the component route that he is beeing asked about.\
To see all current active routes, use:

```
R# show ip eigrp topology active
```

## Stub Routers

A stub router is a router that doesn’t need to know all routes in a network. It is usually the spoke in a hub-and-spoke network. Stub routers do not advertise EIGRP learned routes to other neighbors and Query messages are not sent to stub routers. To define an EIGRP router as stub, use:

```
R(config-router)#eigrp stub [receive-only | [connected] [redistributed] [static] [summary] [leak-map ROUTE-MAP]]
```

By default, if no option is used, IOS will use the **connected** and **summary** options.

* **receive-only** – does not advertise any routes
* **connected** – advertises connected routes but only for interfaces matched with a network command
* **summary** – advertises auto-summarized or statically configured summary routes
* **static** – advertises static routes, assuming the redistribute static command is configured
* **redistributed** – advertises redistributed routes, assuming redistribution is configured
* **leak-map** – allows a subset of the routing table to be advertised. This subset will be matched by the ROUTE-MAP

Usually, the upstream router of the stub router is configured to only send a summary or a default route to the stub router, but this has to be manually configured.

## Next Hop Self

By default, an EIGRP router will advertise routes with itself as the next hop. This can be changed per interface, with:

```
R(confi-if)#no ip next-hop-self eigrp AS-NUMBER
```

The command can be usefull in hub & spoke networks, when you want to have the spokes as the next hop instead of the hub.


# OSPF


# OSPF 101

## Starting the routing process

OSPF can be configured inside the routing process or on the interface. Even a combination of those will be a valid option. The interface configuration will override the routing process configuration.

### Inside the routing process

```
R(config)# router ospf PROCESS
R(config-router)# network NETWORK-ADDR WILDCARD area AREA
```

Depending on the IOS version, when 2 network definitions overlap, the most specific or the last one entered will override the other settings.

### On the interface

```
R(config)# interface INTERFACE
R(config-if)# ip ospf PROCESS area AREA
```

To see what interfaces run OSPF, use:

```
R#sh ip ospf interface brief
Interface    PID   Area            IP Address/Mask    Cost  State Nbrs F/C
Fa0/0        345   0               3.45.0.4/24        10    DR    1/1
Lo0          345   4               4.4.4.4/32         1     LOOP  0/0
Fa0/1        345   4               14.0.0.4/24        10    WAIT  0/0
```

OSPF will advertise a secondary subnet if it is running on the primary subnet. OSPF treats secondary subnets as stub networks so it will not send Hellos on them. Consequently, it will not form adjacencies.

### Split Horizon

OSPF does not use Split Horizon.

### Passive interfaces

When defining a passive interface, the OSPF process will advertise the network but will not send or accept OSPF messages on the interface

```
R(config-router)# passive-interface INTERFACE
```

Or, you can enable passive interfaces by default and then disable it on the interfaces where OSPF should run:

```
R(config-router)# passive-interface default
R(config-router)# no passive-interface INTERFACE
```

## Neighbors

OSPF routers send Hello packets out all OSPF enabled interfaces. When two routers receive each other’s Hellos they can become neighbors and form adjacencies. When the adjacency is up, each router will send its LSAs to the other neighbor. Each router receiving LSAs from a neighbor records the LSA in its LSDB and sends a copy of it to all of its neighbors (LSA Flooding). When the database is complete, each router uses the SPF algorithm to calculate a Loop-Free graph with itself as the root.

### Static Neighbors

Static neighbors can be defined using:

```
R(config-router)# neighbor NEIGH-ADDR [{priority PRI} | {poll-inteval DEAD-TIMER}] 
```

Defining static neighbors is a MUST on non-broadcast and point-to-multipoint non-broadcast networks. Only one side of the connection must be set with the neighbor command. The router that receives unicast hellos will respond with unicast messages to that neighbor.

### Adjacencies

When a router receives a Hello, it will check a list of parameters from the Hello packet against the values configured on the receiving interface. If they do not match, the two routers will not form an adjacency.

* **Area ID**
* **Authentication**
* **Network Mask** (except Point-to-point and virtual link interfaces)
* **HELLO interval, DEAD interval**
* **MTU size** – It can be ignored using:

  ```
  R(config-if)# ip ospf mtu-ignore
  ```
* **Stub flag**
* **Options**

If the values match, the router checks the receiving interface Neighbor Table. If the Router ID that sent the HELLO is in the table, the DEAD interval is reset (meaning – “I already saw this neighbor”). If not, it adds the Router ID that sent the Hello to the table.

Each Hello packet contains a list of Router IDs of the originating router’s neighbors. If the receiving router finds its own Router ID in this list, it knows two-way communication is up and an adjacency can be established.

### Neighbor States

* **Down** – No HELLOs have been heard from this neighbor in the last DEAD interval. HELLOs are not sent to Down neighbors unless they are on NBMA networks (in this case Hellos are sent every Poll interval – default:120sec)
* **Attempt** – Applies only to neighbors in NBMA network, where neighbors are manually configured. A DR-eligible router transitions the state of a neighbor to Attempt when the interface to the neighbor first becomes active or when the router is the DR or the BDR. HELLOs are sent to a neighbor in the Attempt state at HELLO interval instead of POLL interval.
* **Init** – Indicates that a HELLO packet has been seen from the neighbor in the last DEAD interval, but 2-way communication has not yet been established. Starting with this state, the Router ID of the neighbor appears in the Neighbor List field of the HELLOs
* **2-Way** – This state indicates that the router has seen its own Router ID in the Neighbor field of the neighbor’s Hello packet. On multi-access networks, neighbors must be in this state or higher to be DR/BDR-eligibile. Also, receiving a Database Description packet from a neighbor in the INIT state transitions it to the 2-way state. DROthers will remain in 2-Way state with the other DROthers, and will only get to Full state with the DR and BDR.
* **ExStart** – The router and its neighbor establish a master/slave relationship and determine the initial Database Description sequence number. The neighbor with the highest Router ID becomes the master
* **Exchange** – The router sends Database Description packets to describe its entire LSD to neighbors that are in this state. The router may send LSRs to routers in this state, requesting for updated LSAs
* **Loading** – The router sends LSRs to neighbors that are in the Loading state, requesting more recent LSAs that have been discovered in the Exchange state but have not been received yet
* **Full** – The neighbors are fully adjacent and the adjacencies appear in the Router LSA and Network LSA

### Authentication

OSPF supports 3 types of authentication, but actually one of them is null authentication, which means no authentication. The 3 types of authentcation are: Type 0 (Null authentication), Type 1 (Plain Text) and Type 2 (MD5). Null authentication is used by default. To define type 1 or type 2 authentication, use the following commands:

1. Define the authentication type, per area

   ```
   !Per Area:
   R(config-router)# area AREA authentication [message-digest]
   ! <cr> = Type 1 - Plain Text Authentication
   ! message-digest = Type 2 - MD5 Authentication
   ```

   Or per interface:

   ```
   R(config-if)# ip ospf authentication [message-digest|null]
   ! <cr> = Type 1 - Plain Text Authentication
   ! null = Type 0 - No authentication
   ! message-digest = Type 2 - MD5 Authentication
   ```
2. Define the key to be used on the interface:

   ```
   !For plain text:
   R(config-if)# ip ospf authentication-key KEY-STRING
   !For MD5:
   R(config-if)# ip ospf message-digest-key KEY-ID md5 KEY-STIRNG
   !For MD5 both the KEY-ID and the KEY must match
   ```

Both the KEY-ID and the KEY-STRING must match for the routers to form an adjacency. The router will use the last entered key (youngest), but uses an algorithm where it sends multiple copies of the packets, each authenticated with a different key, until all neighbors use the same key. Then it uses the youngest key again.

## Timers

### Hello interval

Hello packets are sent every Hello interval

```
R(config-router)# ip ospf hello-interval HELLO
!10 sec on broadcast networks
!30 sec on non-broadcast networks
```

### Dead interval

If a router does not receive HELLOs from a neighbor in the DEAD interval it will consider the neighbor down:

```
R(config-router)# ip ospf dead-interval DEAD
!Default = 4x HelloInterval
```

When changing the HELLO time, the DEAD interval will be auto-set to 4xHello interval.

#### **Sub second convergence (Fast Hellos)**

Sets the Dead interval to 1 sec and set how many Hellos are sent in that interval

```
R(config-if)#ip ospf dead-interval minimal hello-multiplier MULTIPLIER
```

### Pacing

Pacing is a feature that groups similar packets in a single packet if they have to be sent in a short period of time.\
When LSA flooding occurs, OSPF will wait group toghether LSAs and will send them once every FLOOD-PACING interval:

```
R(config-router)# timers pacing flood MSEC
! Default: 33 msec
```

Evenry LSAs must be refreshed, checksumed or agedout. To avoid a synchronization issue, a router will group together and send packets for any of the previously mentioned operations only once every LSA-GROUP-PACING interval, instead of sending a packet for each LSA.

```
R(config-router)# timers pacing lsa-group SEC
! Default: 240 sec
```

Multiple retransmissions can be grouped into a single packet every RETRANSMISSION-PACING interval

```
R(config-router)# timers pacing retransmission MSEC
! Default: 66 msec
```

### Throttling

Throttling is a feature that rate-limits the OSPF packets.\
When a LSA is first generated, it is sent immediately. The next one is sent after MIN mseconds, and then, for each LSA that is sent, the INCREMENT timer is added. But the total timer between LSAs cannot be more than the MAX timer.

```
R(config-router)# timers throttle lsa all MIN INCREMENT MAX
! Default: MIN = 0 msec, INCREMENT = 5000 msec, MAX = 5000 msec
```

To avoid SPF recalucation in case of link flapping, when a new SPF calculation is required, it is delayed the DELAY interval. If a new calculation is required in the HOLD-TIME, it is delayed and the value of the HOLD-TIME increases with another HOLD-TIME. The total HOLD-TIME value cannot be larger then MAX-HOLD-TIME

```
R(config-router)# timers throttle spf DELAY HOLD-TIME MAX-HOLD-TIME
! Disabled by default.
```

### Other timers

OSPF will accept the same LSA from a neighbor only if the specified time has passed since the last time it received the LSA:

```
R(config-router)# timers lsa arrival MSEC
! Default: 1000 msec
```

If an adjacency must be reset, the router will wait for the maximum value between the RESYNC-TIMEOUT and the interface DEAD interval:

```
R(config-if)# ip ospf resync-timeout SEC
!default: 40 sec
```

If a LSA is not acknowledged, it can be retransmitted after RETRANSMIT-INTERVAL timesout:

```
R(config-if)# ip ospf retransmit-interval SEC
! Default: 5 sec
```

The serialization delay needed to send a LSA out. This value will be added to the LSA age when sending out. It should be greater on low-speed interfaces.

```
R(config-if)# ip ospf transmit-delay SEC
! Default: 1 sec
```

## Packets

OSPF uses IP protocol 89 to send OSPF packets. It sends packet as unicast or multicast to 224.0.0.5 (All OSPF Routers) and 224.0.0.6(All OSPF DRs).

### Packets used in database exchange

Packets sent at this stage are sent as unicast to each router.

* **Database Description** carries a summary description of each LSA. They are used by the receiving router to check if it has the latest LSAs. If a neighbor sees that it doesn’t have a LSA or the latest version of it, it will place the LSA on the Link State Request List
* **Link State Request** is sent to request missing LSAs
* **Link State Update** is sent in response to a LSR with the missing LSA information. As LSAs are received, they are removed from the Link State Request List. All LSAs sent in updates must be acknowledged individually, so before they are sent, they are placed in a Link State Retransmission List and are removed when they are acknowledged.
* **Acknowledgements** can be explicit – A Link State ACk is received containing the LSA Header, or implicit – an update that contains the same instance of the LSA is received

### LSA Flooding

Flooding occurs whenever a change in LSDB happens. This includes a change in an existing LSA or a new LSA. A LSU can contain multiple LSAs. When a LSA is sent to a neighbor a copy of the LSA is added to the Link State Retransmission list. The LSA is retransmitted every Rxmt interval until an ACK is received and it is removed from the Link State Retransmission list. LSUs containing retransmissions are always unicast

Acknoweldgements can be

* **implicit** – A neighhbor implcitly acknowledges an LSA by including a duplicate of the LSA in an LSU back to the originator
* **explicit** – A LSAck is sent back to the originator containing the acknowledged LSA headers
* **direct** – Sent when a duplicate LSA is received from a neighbor (it didn’t receive a previous ACK) or when the LSA’s age is MaxAge and there is no instance of the LSA in the router’s LSDB. They are always sent as unicast
* **delayed** – More LSAs can be acknowledged by a single LSAck. LSAs form multiple neighbors can be acknowledged in a single multicast LSAck. The delayed period must be less than RxmtInterval to avoid retransmission

## Path Selection

The first criteria when choosing OSPF routes is the route type, and then, for the same route type, the metric is considered:

### Route Type

1. **Intra Area (O)** – routes in the same area
2. **Inter Area (O IA)** – routes in another area, but in the same AS
3. **External Type 1 (O E1)** – External paths where the cost is calculated as cost to the ASBR + the cost assigned by the ASBR to the external route
4. **External Type 2 (O E2)** – Ignores the cost to the ASBR when calculating the metric of an external route. It is the default type for external routes!
5. **NSSA Type 1 (O N1)** – external routes from an ASBR in the same NSSA area. Both the cost to the ASBR and the cost assigned by the ASBR are taken into account
6. **NSSA Type 2 (O N2)** – external routes from an ASBR in the same NSSA area, but ignores the cost to the ASBR and considers only the cost assigned by the ASBR to the route

### OSPF Metric

The metric used is the cost of each link in the path.<br>

$$
OSPF\_{metric} = \sum\_{path}Cost
$$

$$
Cost = \frac {BW\_{reference}}{BW\_{interface}}
$$

$$
Cost\_{default} = \frac{10^5}{BW\_{interface}\[kbps]} = \frac{10^8}{BW\_{interface}\[bps]}
$$

\
The reference bandwidth is 100Mbps by default, but can be changed using:

```
R(config-router)# auto-cost reference-bandwidth REFERENCE-BW
! REFERENCE-BW must be entered in Mbps. Default: 100
```

The cost of a link cannot be less than 1 and will be narrowed down to the closes integer. This means that any interface with a bandwidth higher than the reference bandwidth will have the same cost, which is 1.\
The cost can be also manually assigned per interface:

```
R(config-if)# ip ospf cost COST
```

or per neighbor:

```
R(config-if)# ip ospf neighbor cost COST
```

Of course, modifying the bandwidth of an interface will affect the cost:

```
R(config-if)# bandwidth KBPS
```

Using default values, the following costs will be obtained:

| Interface Type     | Default Bandwidth | Default OSPF Cost |
| ------------------ | ----------------- | ----------------- |
| Loopback           | 8.000.000 Kbit    | 1                 |
| Gigabit and higher | 1.000.000 Kbit    | 1                 |
| Fast Etheret       | 100.000 Kbit      | 1                 |
| Ethernet           | 10.000 Kbit       | 10                |
| Serial             | 1544 Kbit         | 64                |

## OSPF Administrative Distance

In OSFP, the default AD for each type of route is 110.\
You can modify the AD filtering on the route type:

```
R(config-router)# distance ospf {[intra-area AD-INTRA] [inter-area AD-INTER] [external AD-EX]}
```

or on the source router (the router that originated the route):

```
R(config-router)# distance AD SOURCE-ROUTER [WILDCARD [ACL]]
```

## Load Balancing

If multiple equal cost paths exist in the final set, OSPF will use them (max 16).

## Filtering

### At the ABR

To filter inter-area routes in or out of an area, use:

```
R(config-router)# area AREA filter-list prefix PREFIX-LIST {in|out}
```

### Filter via summarization

You can also filter Type 3 LSA with area range. When using the not-advertise keyword, the summarry route and its children will not be advertised as a Type 3 LSA => Same effect as “area filter-list out”

```
R(config-router)# area AREA range SUMMARY-IP MASK [not-advertise]
```

### Filter via distribute-lists

Distribute lists can only be used to filter inbound routes and only affects the local routing table, not the LSDB. Due to the Link State nature of OSFP, you cannot filter outbound routes with distribute lists.

```
! Using ACLs
R(config-router)# distribute-list ACL in [INTERFACE]
! Using Prefix-list
R(config-rotuer)# distribute-list prefix PREFIX-LIST in [INTERFCE]
! Using gateway - filter based on the source of the update:
R(config-router)# distribute-list gateway PREFIX-LIST-1 [prefix PREFIX-LIST-2] in [INTERFACE]
! Using route maps:
R(config-router)# distribute-list route-map ROUTE-MAP in [INTERFACE]
```

### Filter using stub area

See [OSPF Areas](https://nyquist.eu/ospf-areas/)

### Filter all LSAs

Used to filter routes only one way, similar to RIP passive-interface.\
On broadcast, non-broadcast and point-to-point networks you can block flooding over specified OSPF interfaces

```
R(config-if)# ip ospf database-filter all out
```

On point-to multipoint networks you can block flooding to a specified neighbor

```
R(config-if)# neighbor NEIGH-ADDR database-filter all out
```

## Summarization

You cannot summarize inside an area because every router must know all paths in an area it belongs to. Summarization can therefore be done only between areas or when redistributing. When advertising summary routes, an OSPF router will auto-generate a discard route (Summary to NULL0). This behavior can be disabled using:

```
R(config-router)# no discard route [internal | external]
! internal - will not generate NULL route for internal summaries
! external - will not generate NULL route for external summaries
```

### Summarization at the ABR

One point in the network where summarization can be configured is at an ABR. The ABR can send only one Type 3 LSA for multiple destinations in the area:

```
R(config-router)# area AREA range SUMMARY-IP MASK [cost COST] [advertise|not-advertise]
!advertise – advertises the summary route using Type 3 LSA
!not-advertise - suppresses the summary route AND the children
```

### Summarization at the ASBR

Another point in the network where summarization can be configured is at an ASBR. The ASBR can send only one Type 7 LSA for multiple destinations that are redistributed into OSPF, using:

```
R(config-router)# summary-address SUMMARY-IP MASK [not-advertise] [nssa-only] [tag TAG]
! not-advertise - suppresses the Type 7 LSA
! nssa-only - the LSAs are not sent outside the NSSA area
! tag - adds a TAG to this route
```

## Default Routes

### Default routes in normal areas

OSPF routers do not advertise default routes by default, but you can change this with:

```
R(config-router)# default-information originate [always] [metric METRIC] [metric-type {1|2}] [route-map ROUTE-MAP]
```

The router will inject a Type 5 LSA advertising a default route if it already has a default route in the routing table.\
Use **always** to generate a Type 5 LSA advertising a default route even when there is no default route in the routing table. If the keyword is not used, the router will never generate a default route unless it has one in the routing table.

Conditional advertising of the route can be achieved using a route-map.

Redistribute static subnets command is used to redistribute static routes in OSPF. However static default route is not injected in to OSPF topology database this way

### Default routes in Stub areas

ABRs advertise default routes into a stub area instead of the Type 4 and 5 LSAs

```
R(config-router)# area AREA default-cost COST
!Assigns a specific cost to the default summary route used for stub areas
```

### Default routes in Totally Stubby Areas

ABRs inject only a Type 3 LSA advertising a default route into a Totally Stubby Area instead of the Type 3, 4 and 5 LSAs

### Default routes in Not So Stubby Areas

ABRs do not advertise default routes into NSSAs by default. To inject a Type 7 LSA advertising a default route into the NSSA, use:

```
R(config-router)# area AREA nssa default-information-orginate
```

### Default routes in NSSA Totally Stubby Areas

ABRs advertise default routes into NSSA Totally Stubby Areas by default as Type 3 LSA


# OSPF Areas

Areas are identified by a 32 bit Area ID. This can be represented as a number in decimal or in dotted decimal format. Area 0 (0.0.0.0) is reserved for backbone. The backbone is responsible for summarizing the topologies of each area to every other area => all inter-area traffic must pass through the backbone.

## Area Types

* **Normal** – Allowed LSAs: 1,2,3,4,5
* **Stub** – Allowed LSAs: 1,2,3 + Default Route as LSA 3 instead of LSAs 4 and 5
* **Totally Stubby** – Allowed LSAs: 1,2 + Default Route as LSA 3 instead of LSAs 3, 4 and 5
* **NSSA** – Allowed LSAs: 1,2,3 + LSA 7 for external routes from the local ASBR
* **NSSA Totally Stubby** – Allowed LSAs: 1,2 + Default route as LSA3 insted of LSAs3,4 and 5 + LSA 7 for external routes from the local ASBR

### Stub Area

* A stub area is an area into which LSA type 4 (ASBR Summary LSA) and 5(AS External LSA) are not flooded.
* As a resut, no external routes will be advertised into the area.
* Instead, ASBRs at the edge of the area use Type 3 LSAs to advertise a default route into the area
* LSAs allowed: 1,2,3 + Default Route instead of LSAs 4 and 5
* Routers configured for stub areas will not form adjacencies with routers not configured for stub ares
* Virtual Links cannot transit a Stub Area
* No router within a stub area can be an ASBR
* A Stub Area can have multiple ABRs but the routers inside cannot determine the best path to an ASBR (ABRs only advertise the default route)

```
R(config-router)# area AREA stub
R(config-router)# area AREA default-cost COST
! Assigns a specific cost to the default summary route used for stub areas
```

### Totally Stubby Area

* Uses a default route to reach AS external routes and inter-area routes
* ABRs only advertise a default route using a Type 3 LSA
* LSAs allowed: 1,2 + Default Route as LSA 3 instead of LSAs 3, 4 and 5(inter-area and as external routes)
* Just the ABRs will have to be configured with the no-summary option, the other routers can be configured only as stub

```
R(config-router)# area AREA stub no-summary
R(config-router)# area AREA default-cost COST
! Assigns a specific cost to the default summary route used for stub areas
```

### Not So Stubby Area (NSSA)

* Stubby Areas with an ASBR attached
* Since Type 5 LSAs are not allowed in Stub areas, the ASBR will originate a type 7 LSA to advertise external routes. This LSA flood stops at the ABR.
* The ABR will translate the Type 7 LSA into a Type 5 LSA to advertise it into area 0. If there are multiple ABRs, only one of them will be elected to translate LSA 7 into LSA 5 – the one with the highest Router ID
* ABR in a NSSA will not generate default routes, unless specified. If it injects a default route, then it will be sent as a Type 7 LSA
* LSAs allowed: 1,2,3,7

```
R(config-router)# area AREA nssa [no-redistribution] [default-information-originate] 
!default-information-origiante = ABR will insert a default route into the NSSA area as LSA type 7
!no-redistribution = ASBR will not insert external routes into the NSSA
```

Normally, NSSA Type 7 routes are redistributed into area 0 as Type 5 LSAs wich points the next hop to the ASBR that introduced the route. You can change this behavior with the following command:

```
R(config-router)# area AREA nssa translate type7 suppress-fa
```

The **translate type7 suppress-fa** keywords on the ABR will force it to translate Type-7 to Type-5 LSAs but change the next-hop to 0.0.0.0 when advertising into area 0. Heaving 0.0.0.0 in the Forward Address of an LSA means to use the advertising router’s address. Otherwise, the Type 5 LSA will have the ASBR’s address in the Forward Address field, which may be unreachable from other areas.

```
ASBRs can also summarize the routes inserted into OSPF with the command:
R(config-router)# summary address PREFIX MASK [not-advertise] [tag TAG]
!used on the ASBR to insert a summary route as LSA 7
```

### NSSA Totally Stubby

* The ABRs use a Type 3 LSA to advertise a default route instead of all other LSA types 3,4 and 5
* The Area also has an ASBR attached that advertises Type 7 LSAs
* ABR in a totally NSSA will generate default information by default into the area as LSA Type 3
* LSAs allowed: 1,2,7 + Default route using LSA 3 from ABR

```
R(config-router)# area AREA nssa no-summary 
```


# OSPF LSAs

## LSA components

* **Sequence number**
  * InitialSequenceNumber = 0x80000001
  * MaxSequenceNumber= 0x7FFFFFFF
  * When MaxSequnceNumber is reached, and a new update must be sent, the LSA is flushed by setting the Age to MaxAge and reflooding it over all adjacencies. Then, a new version of the LSA is sent with sequence number set to InitialSequenceNumber
* **Checksum** – Age is not used in checksum since it’s always changing
* **Age**
  * When a LSA is originated it starts with Age=0. Every time the LSA is flooded out an interface it’s age will be increased with InfTransDelay (1 sec)
  * LSAs are also aged when they reside in the LSDB.
  * When the Age reaches MaxAge (3600 sec = 1 hour), the LSA is reflooded and is flushed from the database
  * When a router needs to flush a LSA it sets the Age to MaxAge

## Link State Database – LSDB

When multiple instances of the same LSA are received, a router will:

1. **Compare sequance number** – the higher will be considered more recent
2. **Compare checksum** – the highest unsigned is considered more recent
3. **Compare Age** – Normally, the highest is considered more recent but if the age differs by more than 15 minutes (MaxAgeDiff) the LSA with the lower Age is considered more recent. If one of the LSAs reached MaxAge it is considered more recent. If none of the previous are true, then the LSAs are considered identical

All valid LSAs received by a router are stored in its LSDB. To view the LSDB, use:

```
R#show ip ospf database
```

Every LSRefreshTime (30min) – the router that originated the LSA floods a new copy of the LSA with an incremented sequence number and an age of zero. Upon receipt, the other OSPF routers replace the old copy of the LSA and begin aging the new copy.

To avoid inefficient LSA updates, LSA group pacing can be enabled. It adds a delay (default 240 sec) to the LSA Update in order to group more LSAs in a single update.(enabled by default)

```
R(config-router)# timers pacing lsa-group SEC
R(config-router)# timers pacing {flood|retransmission} SEC
R# show ip ospf flood-list INTERFACE
!Displays the list of LSAs waiting to be flooded over an interface
```

### Reducing LSA Flooding

LSAs are flooded with DoNotAge=1 and will not be refreshed unless they change. The LSA Age will be incremented as it passes from router to router but it will not age while in the LSDB.

```
R(config-if)#ip ospf flood-reduction
```

## Destination Types

A LSA Destination can be a network or a router.

* If the destination is a **network**, it is added to the Routing Table.
* If the destination is a **router**, it is the route to an ABR or an ASBR. They are kept in an internal table:

  ```
  R# show ip ospf border-routers
  ```

## LSA Types

1. Router LSA
2. Network LSA
3. Network Summary LSA
4. ASBR Summary LSA
5. AS External LSA
6. Group Membership LSA
7. NSSA External LSA
8. External Atribuites LSA
9. Opaque LSA (link-local scope)
10. Opaque LSA (area-local scope)
11. Opaque LSA (AS scope)

To better understand the different types of LSAs, we will use the following example

![LSAs Example Topology](/files/RrTZflLc9u5ABFFVqCed)

Here, we have The backbone area and 3 non-backbone areas (Area 367, Area 1245 and Area 568). Area 568 is configured as an NSSA area where R8 injects a route to 88.88.88.0 as E1. A virtual link exists between R1 and R5. R6 redistributes OSPF into EIGRP while R7 advertises a default route into OSPF and a route to 77.77.77.0 as E2. \
The following show commands will be used in the following segment:

```
R# show ip ospf database
! Lists a summary of all the LSAs in the database
R# show ip ospf database TYPE [LSA-ID]
! TYPE: router|network|summary|asbr-summary|external|nssa-external
! Detailed view of each LSA
R# show ip ospf database TYPE adv-router ROUTER-ID
! Detailed view of LSAs advertised by router ROUTER-ID
R# show ip ospf database TYPE self-originate 
! Detailed view of LSAs advertised by the current router
```

### Router LSA – Type 1

Type 1 LSAs are generated by every router. On each interface running OSPF, the router will generate a Type 1 – Router LSA which will be flooded to all routers in the same area. The LSA ID of a Type 1 LSA is equal to the Router ID that advertised it. For example, on R5:

```
R5#sh ip ospf database

            OSPF Router with ID (5.5.5.5) (Process ID 1)

 Link ID         ADV Router      Age         Seq#       Checksum Link count
1.1.1.1         1.1.1.1         1     (DNA) 0x80000006 0x007315 3
2.2.2.2         2.2.2.2         287   (DNA) 0x80000004 0x00B356 2
3.3.3.3         3.3.3.3         20    (DNA) 0x80000006 0x00C337 2
5.5.5.5         5.5.5.5         510         0x80000003 0x003027 1
! Output omitted
                Router Link States (Area 568)

Link ID         ADV Router      Age         Seq#       Checksum Link count
5.5.5.5         5.5.5.5         261         0x80000009 0x008CA3 1
6.6.6.6         6.6.6.6         268         0x80000002 0x0059D5 1
8.8.8.8         8.8.8.8         388         0x80000002 0x0003EA 2
! Output omitted
                Router Link States (Area 1245)

Link ID         ADV Router      Age         Seq#       Checksum Link count
1.1.1.1         1.1.1.1         1015        0x80000003 0x009944 1
2.2.2.2         2.2.2.2         349         0x80000003 0x00CCE3 1
4.4.4.4         4.4.4.4         290         0x80000005 0x007C3B 5
5.5.5.5         5.5.5.5         1013        0x80000006 0x00DDA3 3
```

On R5 we have 4 Router LSAs in Area 0 (over the Virtual Link to R1), 3 Router LSAs in area 568 and 4 Router LSAs in area 1245. You can see that the Router LSAs are flooded inside the area because R5 has LSAs for R1 and R2 in area 1245 even if they are not directly connected. You can also see that only routers that are inside the area appear in the Router LSAs.\
A Type 1 – Router LSA, contains information about the links a router has in an area. For example, seen from R4, the links of R5 in area 1245 appear like this

```
R5#sh ip ospf database router 4.4.4.4

OSPF Router with ID (5.5.5.5) (Process ID 1)

Router Link States (Area 1245)

LS age: 1502
Options: (No TOS-capability, DC)
LS Type: Router Links
Link State ID: 4.4.4.4
Advertising Router: 4.4.4.4
LS Seq Number: 80000009
Checksum: 0x9325
Length: 84
Number of Links: 5

Link connected to: another Router (point-to-point)
(Link ID) Neighboring Router ID: 5.5.5.5
(Link Data) Router Interface address: 45.45.0.4
Number of TOS metrics: 0
TOS 0 Metrics: 64

Link connected to: a Stub Network
(Link ID) Network/subnet number: 45.45.0.0
(Link Data) Network Mask: 255.255.255.0
Number of TOS metrics: 0
TOS 0 Metrics: 64

Link connected to: a Transit Network
(Link ID) Designated Router address: 14.14.0.1
(Link Data) Router Interface address: 14.14.0.4
Number of TOS metrics: 0
TOS 0 Metrics: 10

Link connected to: a Transit Network
(Link ID) Designated Router address: 24.24.0.2
(Link Data) Router Interface address: 24.24.0.4
Number of TOS metrics: 0
TOS 0 Metrics: 10

Link connected to: a Stub Network
(Link ID) Network/subnet number: 4.4.4.4
(Link Data) Network Mask: 255.255.255.255
Number of TOS metrics: 0
TOS 0 Metrics: 1
```

Notice that this LSA is for area 1245, therefor it has no information about R5’s links in other areas. Also notice the different kinds of links:

* **another router (point-to-point)** – Connection to another router (identified by Router ID) over a point-to-point link. Information about the actual subnet used will be found in the following Stub link. On a point-to-multipoint links, only one stub entry is used for all routers, since they share the same subnet information
* **Stub** – contains subnet information regarding point-to-point or loopback links
* **Transit Network** – contains information regarding the DR on a multi-acesss link. Detailed information about this link can be found in a Type 2 – Network LSA generated by the DR
* **Virtual Link** – contains information regarding the other end of a Virtual Link.

You can see the virtual link in area 0:

```
R1#sh ip ospf database router 5.5.5.5

OSPF Router with ID (1.1.1.1) (Process ID 1)

Router Link States (Area 0)

Routing Bit Set on this LSA
LS age: 1 (DoNotAge)
Options: (No TOS-capability, DC)
LS Type: Router Links
Link State ID: 5.5.5.5
Advertising Router: 5.5.5.5
LS Seq Number: 80000003
Checksum: 0x3027
Length: 36
Area Border Router
AS Boundary Router
Number of Links: 1

Link connected to: a Virtual Link
(Link ID) Neighboring Router ID: 1.1.1.1
(Link Data) Router Interface address: 45.45.0.5
Number of TOS metrics: 0
TOS 0 Metrics: 74
```

### Network LSA – Type 2

Network LSAs are generated by the DR on every multi-access network. The LSA ID is equal to the IP address of the DR that advertised it. This type of LSA is similar to the Router LSA because it is flooded to all routers in an area. You can verify this on R5:

```
R5#sh ip ospf database
                Net Link States (Area 0)

Link ID         ADV Router      Age         Seq#       Checksum
123.123.0.3     3.3.3.3         1262  (DNA) 0x80000001 0x003DDA
!Output Omitted
                Net Link States (Area 568)

Link ID         ADV Router      Age         Seq#       Checksum
56.8.0.5        5.5.5.5         439         0x80000001 0x003357
!Output Omitted
                Net Link States (Area 1245)

Link ID         ADV Router      Age         Seq#       Checksum
14.14.0.4       4.4.4.4         981         0x80000002 0x00589C
24.24.0.4       4.4.4.4         981         0x80000002 0x008F4D

!Output Omitted
```

In each area where R5 is part of, we have a multi-access network. Therefor we have one Network LSA in Area 0 (over the virtual link), one network LSA in area 568(because we used the default NON\_BROADCAST network type), and 2 Network LSAs in area 1245 (because we use ethernet connections on the links from R1 to R4 and R5). This information is flooded to all routers in the area.\
Here’s a detailed look at one Type 2 LSA:

```
R5#sh ip ospf database network 14.14.0.4

            OSPF Router with ID (5.5.5.5) (Process ID 1)

                Net Link States (Area 1245)

  Routing Bit Set on this LSA
  LS age: 1400
  Options: (No TOS-capability, DC)
  LS Type: Network Links
  Link State ID: 14.14.0.4 (address of Designated Router)
  Advertising Router: 4.4.4.4
  LS Seq Number: 80000002
  Checksum: 0x589C
  Length: 32
  Network Mask: /24
        Attached Router: 4.4.4.4
        Attached Router: 1.1.1.1
```

The LSA contains the attached routers along with subnet information (Link State ID + Network Mask). There is no cost information, because this information is found on each of the Type 1 LSAs generated by the Attached Routers.

### Network Summary LSA – Type 3

Network Summary LSAs are originated only by ABRs. An ABR has connections in Area 0 and and at least one non-backbone area. The ABR will advertise intra-area routes of the non-backbone area inside area 0, and will advertise intra-area and inter-area routes from area 0 to the non-backbone areas.\
For example, on R7:

```
R7#sh ip ospf database
!Output omitted
                Summary Net Link States (Area 367)

Link ID         ADV Router      Age         Seq#       Checksum
1.1.1.1         3.3.3.3         1753        0x80000002 0x006DB3
2.2.2.2         3.3.3.3         1753        0x80000002 0x003FDD
3.3.3.3         3.3.3.3         1753        0x80000002 0x00AC76
4.4.4.4         3.3.3.3         1753        0x80000002 0x0047C3
5.5.5.5         3.3.3.3         1753        0x80000002 0x009B2B
8.8.8.8         3.3.3.3         209         0x80000001 0x0095E5
14.14.0.0       3.3.3.3         1756        0x80000002 0x009669
24.24.0.0       3.3.3.3         1756        0x80000002 0x009B50
45.45.0.0       3.3.3.3         1756        0x80000002 0x000F72
56.8.0.0        3.3.3.3         506         0x80000002 0x00BF9B
123.123.0.0     3.3.3.3         1756        0x80000002 0x0082AC
!Output Omitted
```

You can see that the ABR (R3) is advertising Summary LSAs into Area 367 with information regarding all known routes in area 0, except those internal to Area 367.\
In the same time, in Area 0, R3 advertise Summary LSAs with information regarding Intra-Area routes of Area 367:

```
R3#sh ip ospf database | i 3.3.3.3|Link|Summary
!Output Omitted
                Summary Net Link States (Area 0)
Link ID         ADV Router      Age         Seq#       Checksum
6.6.6.6         3.3.3.3         1897        0x80000002 0x008686
7.7.7.7         3.3.3.3         1897        0x80000002 0x0058B0
36.7.0.0        3.3.3.3         1897        0x80000002 0x006793
```

Also in Area 0, R1, R2 and R5 advertise Summary LSAs for their non-backbone areas.

```
R1#sh ip ospf database | e 3.3.3.3
!Output Omitted
                Summary Net Link States (Area 0)

Link ID         ADV Router      Age         Seq#       Checksum
4.4.4.4         1.1.1.1         1970        0x80000002 0x001FFD
4.4.4.4         2.2.2.2         2034        0x80000002 0x000118
4.4.4.4         5.5.5.5         6     (DNA) 0x80000001 0x00C611
5.5.5.5         1.1.1.1         1970        0x80000002 0x007365
5.5.5.5         2.2.2.2         2034        0x80000002 0x00557F
5.5.5.5         5.5.5.5         6     (DNA) 0x80000001 0x0016FD
8.8.8.8         5.5.5.5         1     (DNA) 0x80000001 0x000EB9
14.14.0.0       1.1.1.1         1973        0x80000002 0x006EA3
14.14.0.0       2.2.2.2         2036        0x80000002 0x00B44F
14.14.0.0       5.5.5.5         6     (DNA) 0x80000001 0x007A48
24.24.0.0       1.1.1.1         1973        0x80000002 0x00D71C
24.24.0.0       2.2.2.2         2036        0x80000002 0x0055A4
24.24.0.0       5.5.5.5         6     (DNA) 0x80000001 0x007F2F
45.45.0.0       1.1.1.1         1973        0x80000002 0x00E6AC
45.45.0.0       2.2.2.2         2036        0x80000002 0x00C8C6
45.45.0.0       5.5.5.5         6     (DNA) 0x80000001 0x000C82
56.8.0.0        5.5.5.5         6     (DNA) 0x80000001 0x003A6E
!Output Omitted
```

When an ABR originates a type 3 LSA, it includes the cost from itself to the destination LSA it is advertising. If an ABR knows multiple routes to a destination, it originates a single type 3 LSA with the lowest cost of the multiple routes.\
Here’s how a Type 3 LSA looks like:

```
R7# sh ip ospf data summ 4.4.4.4

            OSPF Router with ID (7.7.7.7) (Process ID 1)

                Summary Net Link States (Area 367)

  Routing Bit Set on this LSA
  LS age: 104
  Options: (No TOS-capability, DC, Upward)
  LS Type: Summary Links(Network)
  Link State ID: 4.4.4.4 (summary Network Number)
  Advertising Router: 3.3.3.3
  LS Seq Number: 80000004
  Checksum: 0x43C5
  Length: 28
  Network Mask: /32
        TOS: 0  Metric: 21
```

When another router receives a type 3 LSA, it will not run SPF and will insert the route in the routing table with a cost equal to the cost included in the LSA plus the cost of the route to the ABR.

```
R7# sh ip route | i 4.4.4.4
O IA    4.4.4.4 [110/31] via 36.7.0.3, 00:20:05, FastEthernet0/1
```

You can see that the route to 4.4.4.4 has a metric of 31 in the routing table, but it was advertised in the LSA with a metric of 21.

### ASBR Summary LSA – Type 4

A Type 4 LSA – ASBR Summary advertises an ASBR, not a subnet. To advertise external destinations, the Type 5 LSAs are used. They advertise external routes with an ASBR as the next hop. This ASBR can be found in the Type 4 LSAs. Each ASBR advertises a Type 4 LSA with its own Router ID as the LSA ID. These LSAs are flooded by the ABRs in all areas (except stub areas). You will find the LSA generated by R7 on R4, beeing advertised by the 2 ABRs in area 1245:

```
R4# show ip ospf database
!Output Omitted
                Summary ASB Link States (Area 1245)

Link ID         ADV Router      Age         Seq#       Checksum
7.7.7.7         1.1.1.1         190         0x80000003 0x00DE27
7.7.7.7         2.2.2.2         224         0x80000003 0x00C041
!Output Omitted
```

Here’s a detailed look at this LSA:

```
R4#sh ip ospf database asbr-summary 7.7.7.7

            OSPF Router with ID (4.4.4.4) (Process ID 1)

                Summary ASB Link States (Area 1245)

  Routing Bit Set on this LSA
  LS age: 502
  Options: (No TOS-capability, DC, Upward)
  LS Type: Summary Links(AS Boundary Router)
  Link State ID: 7.7.7.7 (AS Boundary Router address)
  Advertising Router: 1.1.1.1
  LS Seq Number: 80000003
  Checksum: 0xDE27
  Length: 28
  Network Mask: /0
        TOS: 0  Metric: 20
```

### AS External LSA – Type 5

AS External LSAs are also originated by ASBRs. They advertise a destination external to the OSPF domain or an external default route. They are flooded throughout the AS. The ASBR creates one LSA for each prefix that it injects into OSPF, inluding, if necessary, one for the default route. The LSA-ID is the injected prefix.

```
R4# show ip ospf database
!Output Omitted
                Type-5 AS External Link States

Link ID         ADV Router      Age         Seq#       Checksum Tag
0.0.0.0         7.7.7.7         765         0x80000003 0x006430 1
77.77.77.0      7.7.7.7         5           0x80000001 0x003666 0
79.79.0.0       7.7.7.7         5           0x80000001 0x00568F 0
```

Here’s a detailed view of one of them:

```
R4#sh ip ospf database external 77.77.77.0

            OSPF Router with ID (4.4.4.4) (Process ID 1)

                Type-5 AS External Link States

  Routing Bit Set on this LSA
  LS age: 176
  Options: (No TOS-capability, DC)
  LS Type: AS External Link
  Link State ID: 77.77.77.0 (External Network Number )
  Advertising Router: 7.7.7.7
  LS Seq Number: 80000001
  Checksum: 0x3666
  Length: 36
  Network Mask: /24
        Metric Type: 2 (Larger than any link state path)
        TOS: 0
        Metric: 20
        Forward Address: 0.0.0.0
        External Route Tag: 0
```

The Forward Address contains the next hop address for the destination. Normally, it is set to 0.0.0.0 which means that the Advertising Router should be used as a next hop. The exception occurs when a Type 7 LSA is translated to a Type 5 LSA, which by default sets the Next Hop Address to the ASBR that first advertised the prefix.\
Also notice the Metric Type information. When the metric type is set to 2, then the router will use the metric in the LSA to compute the route metric:

```
R4#sh ip route | i 77.77.
O E2    77.77.77.0 [110/20] via 24.24.0.2, 01:01:40, FastEthernet0/0
```

When the metric type is set to 1, the router will add the metric to the ASBR to the metric in the Type 5 LSA.

### NSSA External LSA – Type 7

NSSA External LSAs are originated by ASBRs in NSSAs. They are flooded only within the NSSA area in which it was originated. The ABRs will translate to Type 5 AS Summary LSA before flooding to area 0. Such a LSA can be seen on R5:

```
R5# show ip ospf database
!Output Omitted
                Type-7 AS External Link States (Area 568)

Link ID         ADV Router      Age         Seq#       Checksum Tag
88.88.88.0      8.8.8.8         548         0x80000002 0x0085C6 0
```

Type 7 LSAs differ from Type 5 LSAs in that they set the Next Hop to the ASBR address instead of 0.0.0.0 – which means use the Advertising Router’s address. This will propagate by default into the translate Type 5 LSAs that go into area 0. We can observe this by looking first at the Type 7 LSA:

```
R5#sh ip ospf database nssa-external 88.88.88.0

            OSPF Router with ID (5.5.5.5) (Process ID 1)

                Type-7 AS External Link States (Area 568)

  Routing Bit Set on this LSA
  LS age: 208
  Options: (No TOS-capability, Type 7/5 translation, DC)
  LS Type: AS External Link
  Link State ID: 88.88.88.0 (External Network Number )
  Advertising Router: 8.8.8.8
  LS Seq Number: 80000004
  Checksum: 0xFDCC
  Length: 36
  Network Mask: /24
        Metric Type: 1 (Comparable directly to link state metric)
        TOS: 0
        Metric: 20
        Forward Address: 8.8.8.8
        External Route Tag: 0
```

Here, we have the Forward address set to the ASBR ID. When the ABR translates it into a Type 5 LSA, it will have the same forward address:

```
R4#sh ip ospf database external 88.88.88.0

            OSPF Router with ID (4.4.4.4) (Process ID 1)

                Type-5 AS External Link States

  Routing Bit Set on this LSA
  LS age: 349
  Options: (No TOS-capability, DC)
  LS Type: AS External Link
  Link State ID: 88.88.88.0 (External Network Number )
  Advertising Router: 5.5.5.5
  LS Seq Number: 80000004
  Checksum: 0xECF3
  Length: 36
  Network Mask: /24
        Metric Type: 1 (Comparable directly to link state metric)
        TOS: 0
        Metric: 20
        Forward Address: 8.8.8.8
        External Route Tag: 0
```

This behavior can be changed if we use:

```
R(config-router)# area AREA nssa translate type7 suppress-fa
```

This will suppress the ASBR address in the Forward Address field and replace it with 0.0.0.0, meaning “use the advertising router”. This can be useful in situation where the ASBR cannot be reached from outside its area.\
Also notice that this route was injected with a metric type of 1. This means that in the routing table, the metric in this LSA will be added to the metric to the ASBR. Here’s the proof: the route has a metric of 149. 20 represent the LSA and 129 represent the distance to the ASBR.

```
R4#sh ip route | i 88.88
O E1    88.88.88.0 [110/149] via 45.45.0.5, 00:08:55, Serial1/0
R4#sh ip route | i 8.8.8
O IA    8.8.8.8 [110/129] via 45.45.0.5, 01:09:54, Serial1/0
```


# OSPF Mechanics

## OSPF Router ID

Each router selects an OSPF Router ID when the OSPF process starts. The Router ID is a 32 bit number, usually written in dotted decimal format. The selection process is:

1. Manually configured Router ID

   ```
   R(config-router)# router-id ROUTER-ID
   ```
2. Highest Loopback IP address
3. Highest non-Loopback “up/up” IP address

The interface used for Router ID doesn’t have to run OSPF and the router ID chosen when OSPF starts will remain even if the interface changes status or is deleted.\
To use a new Router ID, type:

```
R# clear ip ospf process
```

## OSFP Network Types

| Feature                   | Broadcast                     | Non-Broadcast      | Point-to-point             | Point-to-multipoint | Point-to-multipoint non-broadcast |
| ------------------------- | ----------------------------- | ------------------ | -------------------------- | ------------------- | --------------------------------- |
| Default for               | Ethernet                      | FR physical, FR MP | FR P2P, PPP, HDLC, Tunnels | –                   | –                                 |
| DR?                       | YES                           | YES                | NO                         | NO                  | NO                                |
| Hello Timer               | 10 sec                        | 30 sec             | 10 sec                     | 30 sec              | 30 sec                            |
| Dead Timer                | 40 sec                        | 120 sec            | 40 sec                     | 120 sec             | 120 sec                           |
| Hellos to                 | 224.0.0.5                     | Unicast            | 224.0.0.5                  | 224.0.0.5           | Unicast                           |
| Other packets to          | 224.0.0.5 (B)DR 224.0.0.6 DRO | Unicast            | 224.0.0.5                  | 224.0.0.5           | Unicast                           |
| Static Neighbor           | CAN                           | MUST               | NO                         | CAN                 | MUST                              |
| Multiple adjacencies      | YES                           | YES                | NO                         | YES                 | YES                               |
| Next hop for same segment | Advertising Router            | Advertising Router | Self                       | Self                | Self                              |
| Next hop for diff segment | Self                          | Self               | Self                       | Self                | Self                              |

* 224.0.0.5 is the All OSPF routers multicast address
* 224.0.0.6 is the OSFP DRs mulsticast address

In addition to these network types there is the stub network type that is the default for loopback interfaces. The stub networks will be advertised as /32 routes regardless of the interface mask. To advertise the network according to the mask, the network type should be changed to point-to-point.\
The network type is independent of the interface type, and the default values can be changed with:

```
R(config-if)# ip ospf network {broadcast|non-broadcast|point-to-point|point-to-multipoint [non-broadcast]}
```

To verify the network type, use:

```
R# show ip ospf interface INTERFACE
```

See this article about [running OSPF over Frame Relay](https://nyquist.eu/routing-over-frame-relay/#23_OSPF).\
Routers can become neighbors even if the network type is different, as long as they agree on 2 things: If a DR is required or not, and if the HELLO/DEAD timers are the same. You can’t change weather a DR is required or not, but you can change the timers with the following commands:

```
R(config-if)# ip ospf hello-interval SEC
R(config-if)# ip ospf dead-interval SEC
```

To summarize, here’s what you need to remember:

| Keyword                                   | What it means                                                                | Network Type                                                                                           |
| ----------------------------------------- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
| If it contains the word **POINT**         | <ul><li>No DR Election</li><li>Changes the next hop to self</li></ul>        | <ul><li>Point-to-point</li><li>Point-to-multipoint</li><li>Point-to-multipoint non-broadcast</li></ul> |
| If it contains the word **NON-BROADCAST** | <ul><li>Sends packets as unicast</li><li>Requires static neighbors</li></ul> | <ul><li>Non-Braodcast</li><li>Point-to-multipoint non-broadcast</li></ul>                              |
| If it is **usually used for Frame Relay** | <ul><li>Slow Timers</li></ul>                                                | <ul><li>Non-Braodcast</li><li>Point-to-multipoint</li><li>Point-to-multipoint non-broadcast</li></ul>  |

## Designated Router (DR) Election

When an OSPF router becomes active, it checks for an active DR and BDR on the networks that have a type that requires such a process (broadcast and non-broadcast). The router will wait for *WaitTimer* (=RouterDead Interval) for a DR and BDR to be advertised in a Hello packet before starting an election process.

* If a DR and BDR exist, the router accepts them
* If there is no DR, but there is a BDR, the BDR becomes the DR and an election for BDR is held
* If there is no BDR and no DR, an election is held for both DR and BDR

If an election takes place, this is how the DR is chosen:

* The router with the highest priority becomes the DR/BDR (depending on the election type)

  ```
  R(config-if)# ip ospf priority PRI
  !Default: 1
  !0 = never become a DR
  ```
* In case of a tie, the next criteria is highest Router ID

Since an existing DR and BDR is accepted, the first 2 routers that initialize on a Broadcast network will be selected as DR and BDR.

## Interface Status

* **Down**
* **Point-to-point** – The router sends Hellos and will attempt to establish an adjacency with the other end of the link. Option is available for Point-to-point, point-to-multipoint and virtual links
* **Waiting** – The router sends Hellos and waits for a Hello from DR and BDR. Option is available for Broadcast and NBMA networks
* **DR** – The router is the DR and will establish adjacencies with the DROthers
* **Backup** – The router is the BDR and will establish adjacencies with the DROthers
* **DROther** – The router is neither DR nor BDR and will establish adjacencies only with the DR and BDR, but will send Hellos to all routers
* **Loopback** – The interface is still advertised in Router LSAs even though packets cannot transit such an interfac

## Router Types

* **Internal** – All interfaces are in the same area
* **ABR (Area Border Router)** – Has at least one interface in area 0 and one in another area
* **Backbone Routers** – Routers with at least one interface in area 0
* **ASBR (AS Boundary Router** – Gateways for external trafic, injecting routes form other protocols into OSPF

## Virtual Links

Normally, traffic from one area to another must pass through Area 0. Sometimes, this is physically impossible, so the concept of virtual links was added. A virtual link can be used to create a neighbor adjacency for 2 routers that are not normally neighbors.

A virtual link extends area 0 from one ABR to another router in one of it’s non-zero areas. Therefore, the virtual link can only transit one area. A new adjacency will be formed over this virtual link so the area 0 extends to the other router, making it an ABR. The concept can be extended now, and a new virtual link can be created from this router to another router. Of course, the same rules apply.

Virtual Links cannot transit any flavor of stub areas.

To define a Virtual Link use the following command on the 2 ends of the virtual link. Of course, each router must reference the other router’s Router ID:

```
R(config-router)# area AREA virtual-link ROUTER-ID
```

If area 0 is set for authentication, then the virtual links must also be configured for authentication:

```
R(config-router)# area AREA virtual-link ROUTER authentication [message-digest]
! sets authentication as cleat text or MD5
R(config-router)# area AREA virtual-link ROUTER authentication-key CLEARTEXT-KEY
! sets clear text key
R(config-router)# area AREA_ID virtual-link ROUTER-ID message-digest-key KEY-ID md5 MD5-KEY
! sets md5 key
```

## OSPF over Demand Circuits

Periodic Hellos are suppresed and periodic refreshes of LSAs are not flooded. The circuit is used only at the initial db sync and only when changes have occured, in order to send the updated LSAs. Hellos are still sent over multi-access networks, but are not sent on point to multipoint network types.

```
R(config-if)# ip ospf-demand-circuit
```

Only one end of the Poin-to-pont connection or the multipoint in a Point-to-multipoint need this setting

## OSPF DNS Lookups

By default OSPF will not perform DNS lookup to translate the neighbor IP addresses to their hostnames. To enable, use:

```
R(config)#ip ospf name-lookup
```


# IS-IS


# IS-IS 101

## Starting the routing process

Starting IS-IS process requires a 2 step configuration:\
1\. In the global config

```
R(config)# router isis [AREA-TAG]
!AREA-TAGs are used to run multiple IS-IS processes. Default: NULL
R(config-router)# net NETWORK-ENTITY-TITLE
!NETWORK-ENTITY-TITLE is in NSAP format. E.g 49.0001.0010.0100.1001.00
```

2\. On the interfaces that will be enabled for IS-IS

```
R(config)# interface INTERFACE
R(config-if)# ip router isis
```

## Passive interface

The passive interface command in IS-IS has a basically an opposite meaning to what it means in the other routing protocols. In IS-IS, since you have to manually select the interfaces that will run IS-IS and will send packets to form adjacencies, if you don’t want to run IS-IS on an interface you shouldn’t enable IS-IS on it. But what if you want to advertise it, without making any adjacency on it? Then, it’s a passive interface so go ahead and configure it as such 🙂

```
R(config-router)# passive-interface {INTERFACE-ID|default}
default: all local interfaces will be advertised
```

### Router Levels

By default Cisco IS-IS routers run at both Level1 and Level2. You can change the level on a per interface basis, using the command:

```
R(config-if)# isis circuit-type [level-1|level-1-2|level-2-only]
!Default: level-1-2
```

## Neighbors

Adjacencies are formed through the exchange of HELLO packets. On broadcast interfaces, there are separate HELLOs for each level, but on Point-to-Point interfaces, there is a single L1L2HELLO for efficiency.

### Authentication

Since HELLOs (ILH) are exchanged between neighbors and are not forwarded to other devices and since the packets describing the routes are forwarded to other ISes in the area or domain, there is a different mechanism of authentication for each type of packet:

#### **ILH Authentication**

Authentication for ILHs is done at the interface level:

```
R(config-if)# isis authentication mode {text|md5} [level-1|level2]
! if level is not selected, it applies to both levels.
R(config-if)# isis authentication key-chain KEY-CHAIN [level-1|level-2]
```

On older implementations, use:

```
R(config-if)# isis password PASSWORD
```

#### **LSP, CSNP, PSNP Authentication**

Authentication for these packets needs to be the same in the entire area, so this is done inside the routing process configuration:

```
R(config-router)# authentication mode {text|md5} [level-1|level-2]
! if level is not selected, it applies to both levels.
R(config-router)# authentication key-chain KEY-CHAIN [level-1|level-2]
```

On older implementations, use:

```
R(config-router)# area-password PASSWORD
! applies to Level1
R(config-router)# domain-password PASSWORD
! applies to Level2
```

## Timers

HELLO packets are sent every HELLO-INTERVAL – default 10 sec. The timeout value is based on the HELLO-INTERVAL and a HELLO-MULTIPLIER – default 10×3 = 30 sec. To change thse values, use:

```
R(config-if)# isis hello-interval SEC [level]
! default SEC = 10
R(config-if)# isis hello-multiplier N [level]
! default N = 3
```

However, once an IS is selected as the DIS for an area, it sends Hellos 3 times as fast (10/3 seconds) and has a similar hold value (30/3=10 sec) in order to detect failed DIS quicker.

## Packets

Hello Packets are exchanged every HELLO-INTERVAL (default: 10 sec) in order to create and maintain adjacencies and electing a DIS (similar to OSPF DR). On broadcast interfaces, separate HELLOs (IIH = IS-IS HELLOs) are sent for each level, while on point-to-point interfaces a single IIH is sent for both levels.\
LSP (Link State PDUs) are used to advertise the routing information. LSPs have variable size and include routing information as TLV (type, lenght, value) records inside the LSP.\
CSNP (Complete Sequence Numbers PDU) and PSNP (Partial Sequence Numbers PDU) are packets used to synchronize Link-state database.

## Metric

IS-IS can support several types of metrics, but only the “Default” metric is required to be implemented. It is normally associated with the circuit bandwidth. If other metrics are supported (Delay, Expense, Error) then a new SPF tree is created for each of them. Cisco routers usually support only the Default metric, but each circuit (interface) has a metric of 10, regardless of the bandwidth of that link. It’s up to the admin to change the metrics on the interface with the command:

```
R(config-if)# isis metric DEFAULT-METRIC [DELAY-METRIC] [EXPENSE-METRIC] [ERROR-METRIC] [level-1|level-2]
! if no level is specified, it applies to all levels.
```

Newer versions of the IOS use the Wide Metrics where values can are stored on 24 bits for individual metrics and on 32 bits for cumulative metrics. On older IOS versions, individual metrics had values between 1-63 (6 bits) and cumulative metrics had values between 1-1023 (10 bits).

## Adminsitrative Distance

IS-IS has an AD of 115.


# IS-IS Mechanics – CLNP

## ISO OSI Terminology

| ISO OSI term        | TCP/IP Equivalent |
| ------------------- | ----------------- |
| End System          | Host              |
| Intermediate System | Router            |
| Circuit             | Interface         |
| Area                | Area              |
| Domain              | Autonomous System |

**IS-IS** = Intermediate System to Intermediate System\
**CLNP** = Connection-Less Network Protocol = Layer 3 network protocol that is used to communicate between ESes. CLNP offers a CLNS (Connection-Less Network Service). CLNP uses **NSAP** (Network Service Access Point) addressing. An NSAP address is assigned to an entire node, not to individual interfaces.\
**SNPA** = Sub-Network Point of Attachment = Layer 2 address\
**Local Circuit ID** = When added to the ISIS protocol, each interface gets Local Circuit ID. Values: 1-255.

### NSAP Format

The NSAP address has 2 parts, which can be split into smaller pieces:

* **IDP** = Initial Domain Part
  * **AFI** = Authority and Format Identifier = 1 octet long (00-FF). It identifies the format of the rest of the address. Usually 49 is used, which means the format/authority is local – similar to private addressing
  * **IFI** = Initial Domain Identifier = variable length, can miss if used only inside a domain
* **DSP** = Domain Specific Part
  * **HO-DSP** = High Order Domain Specific Part = variable length,identifies an area of the domain
  * **System ID** = Identifies the node in an area. Length can be between 1 and 8 Bytes. Usually 6 Bytes are used
  * **SEL** = Identifies the service that should process the packets on the node. If SEL=0, the frame is for the node itself and the NSAP address with SEL=0 is called **NET** = Network Entity Title

```
R(config-router)# net 49.0001.1111.1111.1111.00
! 49 = AFI
! 0001 = HO-DSP (Area)
! 1111.1111.1111.1111 = System ID
! 00 = SEL
```

## Routing Levels

**Level 0** = ES to ES on the same link (or ES to IS on the same link). ES and IS send Hellos – ESH and ISH – to advertise each other’s presence on the network.\
**Level 1** = ES to ES in a single area = IS nodes collect lists of ES nodes directly attached to them and exchange this information with all other IS in the area.\
**Level 2** = ES to ES in different areas from the same domain = IS nodes exchange area prefixes\
**Level 3** = ES to ES in different domains = Initially IDRP (Inter Domain Routing Protocol) was used to exchange domain information, but now BGP can be used since now it’s protocol independent.


# BGP


# BGP 101

## Starting the Routing Process

### Define the routing process

Only one BGP process can run on a router and it can be started using:

```
R(config)# router bgp AS-NUMBER
```

The AS Number used to be a 16 bits number ranging from 0 to 65535. According to RFC 4893, the AS number can have 32 bits starting from 0.0 to 65535.65535.

* **0** – reserved
* **1 – 64495** – assignable by IANA for public use
* **64496 – 64511** – reserved for documentation
* **64512 – 65534** – assigned for private use
* **65535 -** reserved

For details, see [IANA](https://www.iana.org/assignments/as-numbers/as-numbers.xml).

### Specify neighbors manually

BGP does not have a method to autodetect neighbors. Instead, each neighbor must be manually defined. See [Neighbors](broken://pages/mPoCpeQgE09PzI5oYsTQ).

### Route advertising

By default, BGP will not advertise any route to its BGP peers. You will have to either redistribute a route from another protocol (redistribute command) or statically define a route that should be advertised (network command). BGP will only advertise routes that already exist in the routing table.

#### **Redistribution**

```
R(config-router)# redistribute ... 
```

#### **Static network definition**

```
R(config-router)# network NETWORK-ADDRESS [mask NETMASK]
! If NETMASK is missing, the default mask of the classful address is used
```

The exact NETWORK-ADDRESS/NETMASK must exist in the routing table, or the route will not be advertised!

#### **Inject Maps**

You can inject routes that do not exist in the routing table using an inject map. A less specific aggregate route must exist in the routing table.

```
R(config-router)# bgp inject-map INJECT-MAP exist-map EXIST-MAP [copy-attributes]
! copy-attributes will copy attributes from the aggregate route to the injected routes
```

The INJECT-MAP will set the prefixes that are to be inserted.

```
R(config)# route-map INJECT-MAP
R(config-route-map)# set ip address prefix-list INJECTED-ADDRESS
```

The EXIST-MAP will match the aggregate route and the neighbor that advertised it to us:

```
R(config)# route-map EXIST-MAP
R(config-route-map)# match ip address prefix-list AGGREGATE-ROUTE
R(config-route-map)#  match ip route-source prefix-list ROUTE-SOURCE
! Both match lines must be configured!
```

## Neighbors

BGP does not have a method to autodetect neighbors. Instead, each neighbor must be manually defined:

```
R(config-router)# neighbor NEIGH-ADDR remote-as NEIGH-AS-NUMBER
```

Each router will attempt to create a TCP connection over port 179 with the neighbor. Depending on the NEIGH-AS-NUMBER, the peering will be considered internal or external and will act differently. BGP also offers the possibility to temporarily disable a neighbor by using the following command:

```
R(config-router)# neighbor NEIGH-ADDR shutdown
```

### Peer Groups and Peer Templates

You can also create a group of neighbors and apply the same setting to all of them. First define the group and configure it with a remote AS:

```
R(config-router)# neighbor PEER-GROUP peer-group
R(config-router)# neighbor PEER-GROUP remote-as AS-NUMBER
```

Make other configurations to the PEER-GROUP and then define neighbors as members of the PEER-GROUP:

```
R(config-router)# neighbor NEIGH-ADDR peer-group PEER-GROUP
```

Another option is to use a Peer Template. There are 2 types of templates: session and policy templates. Session templates define parameters that make the 2 peers to become neighbors.\
First define the template:

```
R(config-router)# template peer-session PEER-TEMPLATE
R(config-router-stmp)# remote-as AS-NUMBER
! You can inherit another template and/or define other parameters
R(config-router-stmp)# inherit peer-session OTHER-TEMPLATE
R(config-router-stmp)# ...
```

In the same way, you can define and a policy template, only this template is used for policy definition for filtering, attribute modifications and so on:

```
R(config-router)# template peer-policy POLICY-TEMPLATE
R(config-router-ptmp)# ...
```

Then apply the template to the neighbor:

```
R(config-router)# neighbor NEIGH-ADDR inherit peer-session PEER-TEMPLATE
R(config-router)# neighbor NEIGH-ADDR inherit peer-policy POLICY-TEMPLATE
```

### iBGP Peers

When the neighbors are in the same AS we have an internal BGP peering. BGP packets for iBGP peers will be sent with a TTL of 255 because BGP assumes the internal BGP peer is not directly connected, but a route to it exists in the routing table.\
It is important to remember the following rules regarding iBGP peers:

* **routes advertised to iBGP peers will not have the next hop modified**. You can use next-hop-self to change this setting:

  ```
  R(config)# neighbor NEIGH-ADDR next-hop-self
  ```
* **routes advertised to iBGP peers will not have an AS\_PATH prepended** since we are advertising inside the same AS.
* **routes learned from iBGP peers will not be advertised to other iBGP peers** to avoid loops. This is known as BGP Split Horizon rule. AS\_PATH is used as a loop avoidance mechanism, and since all routers are in the same AS, we cannot use it inside the AS. BGP was designed considering that inside an AS, all BGP peers are fully meshed. In a fully meshed environment, all iBGP peers would exchange external BGP routes and they would have a similar view of the outside world. However, it is hard to satisfy this requirement in large networks, so there are workarounds: Route Reflectors and Confederations

#### **Route Reflectors**

By default, routes learned via iBGP peers are not advertised to other iBGP peers. By setting a neighbor as a route reflector client, this rule is changed like this:

* If the iBGP route is received from a non route-reflector client peer, reflect the route only to route-reflector clients
* If the iBGP route is received from a route-reflector client, reflect the route to all iBGP peers (router-reflector clients and non route-reflector clients)

It is as if a Route Reflector treats its clients as eBGP peers, but doesn’t do any AS\_PATH prepending and doesn’t change the NEXT\_HOP.

```
R(config-router)# neighbor NEIGH-ADDR route-reflector client
```

#### **Confederations**

A confederation is a group of AS-es that appear to the exterior as a single AS. They will use the Confederation Identifier when peering with other routers that are not part of the same confederation and the locally defined AS number to peer with other members of the confederation. BGP updates learned from other sub-AS-es of a confederation are enclosed between () in the AS\_PATH.

```
R(config)# router bgp SUB-AS-NUMBER
R(config-router)# bgp confederation identifier AS-NUMBER
R(config-router)# bgp confederation peers NEIGH-SUB-AS-NUMBER
! NEIGH-SUB-AS-NUMBER should list the other sub-AS-es of a confederation
```

Prefixes are passed between Sub-AS-es based on eBGP rules, but most attributes are left unchanged, like in iBGP, including the NEXT\_HOP.\
Inside a Sub-AS, the peers should be fully meshed or use Route Reflectors.

#### **Synchronization Rule**

The synchronization rule applies only to iBGP updates and is used to prevent black holes – that is forwarding traffic to a BGP speaker via a non-BGP speaker that will drop all traffic for destinations known only via BGP.\
When synchronization is on, a router will not advertise routes, unless it has the same route in its IGP routing table. To overcome this rule, you can:

1. Enable synchronization and Redistribute BGP in IGP
2. Enable synchronization but run MPLS inside the AS – actually hiding the the packet destination and forwarding based on the MPLS label, thus preventing the blackhole
3. Disable synchronization (default) if you know how to prevent the blackhole:

   ```
   R(config-router)# no synchronization
   ```

When using synchronization with redistribution, the BGP route will usually be marked with **r** – meaning RIB-failure. This happens because, by default, the IGP route has better AD than the BGP route. Still, the route will be considered valid and it will be advertised. To disable advertisement of RIB-failed routes, use:

```
R(config-router)# bgp suppress-inactive
```

The BGP Synchronization rule is independent from other BGP rules, like not advertising internal routes to other internal neighbors.\
When the IGP is OSPF, the route is considered synchronized only if it is received via BGP and OSPF from the same router. This means that the Router ID of OSPF and BGP must match on the advertising router.

### eBGP peers

When the neighbors are in different AS-es we have an external BGP peering. BGP packets for eBGP peers are sent with a TTL of 1 by default, because BGP assumes eBGP peers are directly connected. Problems arise when the peers are not directly connected, and another router is between them. By default, the routers will not form a relationship because their packets don’t reach the other router. To solve this problem, you must change the TTL value of the BGP packets, using:

```
R(config-router)# neighbor NEIGH-ADDR ebgp-multihop [HOP-COUNT]
! Default HOP-COUNT: 255
```

The other option is to use the ttl-security feature:

```
R(config-router)# neighbor NEIGH-ADDR ttl-security hops MIN-HOPS
```

This time, the router will send eBGP packets with a TTL of 255, but will only accept packets from the neighbor if they have a TTL equal to or higher than the value defined in MIN-HOPS. This means that if the neighboring router is N hops away, you must set MIN-HOPS to a value equal to 255-N. You cannot configure both ebgp-multihop and ttl-security for the same neighbor.

The third option that can be used, is a special case where you use the Loopback addresses of each router as the source, but you still want to make them connect over a direct connection. Normally, a router will check if the eBGP peer is directly connected, and if not, it will give up on sending any packet, since it knows it’s TTL will expire. However, you can disable this check and send BGP packets with TTL=1 sourced and destined to the loopback interfaces, using the following steps:

```
! 1. set the loopback as the update source interface:
R(config-router)# neighbor NEIGH-ADDR update-source INTERFACE
! 2. disable connected check
R(config-router)# neighbor NEIGH-ADDR disable-connected-check
```

This configuration will allow peering between the loopbacks but will only use a directly connected network for sending the updates. A non-directly connected network will make the TTL expire in the BGP packets.

For eBGP advertisements we must remember the followings:

* **routes advertised to eBGP peers will have the next-hop set to the IP of the update’s source interface**. The update’s source interface can be manually configured as below, or it is the interface used to send messages to the neighbor.

```
R(config)# neighbor NEIGH-ADDR update-source INTERFACE
```

* **routes advertised to eBGP peers will have the AS\_PATH prepended with the AS of the sending router**
* **routes learned from an eBGP are advertised to both iBGP and eBGP peers. Routes learned from iBGP peers are advertised only to eBGP peers**

### BGP Peering State Machine

* **IDLE** – Waiting to start TCP connection
* **CONNECT** – Waiting to complete TCP connection
* **ACTIVE** -TCP Connection failed. Trying again
* **OPEN SENT** – TCP Connection completed. OPEN message was sent
* **OPEN CONFIRM** – OPEN message received, parameters agree upon, waiting for a keepalive
* **ESTABLISHED** – Peering complete

#### 2.5 Authentication

Authentication uses MD5 hashes and must be configured on each neighbor:

```
R(config-router)# neighbor NEIGH-ADDR password PASS
```

## Timers

### Keepalive and Holdtime

By default, the keepalive timer is 60 seconds, and the hold-time timer is 180 seconds.When a connection is started, BGP will negotiate the hold time with the neighbor. The smaller of the two hold times will be chosen. The keepalive timer is then set based on the negotiated hold time and the configured keepalive time.

```
R(config-router)# timers bgp KEEPALIVE HOLDTIME [MIN-HOLDTIME]
! MIN-HOLDTIME = MIN Holdtime from neighbor
! Default values: KEEPALIVE: 60 sec, HOLDTIME: 180 sec, MIN-HODLTIME: 0 sec
```

The timers can also be set for each neighbor:

```
R(config-router)# neighbor NEIGH-ADDR timers KEEPALIVE HOLDTIME [MIN-HOLDTIME]
```

### Advertisement interval

BGP doesn’t send updates at regular intervals, instead it only sends triggered updates. But in order to prevent unnecessary propagation of link flaps, it will delay sending a new update until the Advertisement Timer expires. This timer starts as soon as the last update finished and it has a value of 5 seconds for iBGP peers and 30 seconds for eBGP peers.\
To modify it, use:

```
R(config-router)# neighbor NEIGH-ADDR advertisement-interval SEC
! Default SEC: 5 sec iBGP, 30 sec eBGP
```

## Packets

* **OPEN** – contains:
  * **Local AS Number**
  * **Local Router ID**
    1. Defined ID:

       ```
       R(config-router)# bgp router-id ROUTER-ID
       ```
    2. Highest IP of a loopback interface
    3. Highest IP of an up/up physical interface
  * **Hold Time** – negotiated to the lowest value
  * **Other options**
* **KEEPALIVE** – By default sent every HoldTime/3 seconds
* **UPDATE**– used to advertise or withdraw a prefix. Includes:
  * Withdrawn routes – routes that should be discared
  * NLRI (Network Layer Reachability Infromation) – routes being advertised
  * Path attributes
* **NOTIFICATION** – sent when an error occurs that causes BGP connection to close

## Metric

BGP uses router attributes to make routing decisions, unlike an IGP that uses link cost. However, when redistributing into BGP, the value of the IGP metric is copied into the [MED attribute](/ipv4/ipv4-routing/bgp/bgp-attributes#route-attributes).

## Administrative Distance

Default iBGP: 20\
Default eBGP: 200

```
R(config-router)# distance AD SOURCE-IP WILDCARD ACL
R(config)# distance bgp EBGP-AD IBGP-AD LOCAL-AD
```

## Filtering

### ACL

```
R(config-router)# neighbor NEIGH-ADDR distribute-list ACL
```

The ACL can be a standard ACL, in which case it will match on prefix bits but with any prefix length, or an extended ACL where it will match on both the prefix bits and the prefix length:

```
R(config)# access-list ACL {permit|deny} PREFIX PREFIX-WILDCARD MASK MASK-WILDCARD
```

Also, the distribute list can be applied to all the BGP process:

```
R(config-router)# distribute-list ACL {in|out} [INTERFACE]
```

### Prefix List

```
R(config-router)# neighbor NEIGH-ADDR prefix-list PREFIX-LIST
```

A prefix list contains a list of permit or deny entires and ends with an impicit deny any.\
To define a prefix-list, use:

```
R(config)# ip prefix-list PREFIX-LIST NETWORK/LEN [ge MIN] [le MAX]
```

1. If only the NETWORK/LEN appears in the prefix entry, it will only match this prefix.
2. If the **ge** keyword is used, it will match all prefixes that have the first LEN bits the same as NETWORK and that have a subnet mask length between MIN and 32.
3. If the **le** keyword is used, it will match all prefixes that have the first LEN bits the same as NETWORK and that have a subnet mask length between 0 and MAX.
4. If both **le** and **ge** are used, it will match all prefixes that have the first LEN bits the same as NETWORK and that have a subnet mask length between MIN and MAX. Also, only entries where LEN\<MIN<=MAX are valid.

Here are some examples:

```
R(config)# ip prefix-list EXAMPLE permit 0.0.0.0/0
! Matches only the default route
R(config)# ip prefix-list EXAMPLE 0.0.0.0/0 le 32
! Matches any route
R(config)# ip prefix-list EXAMPLE 0.0.0.0/0 {le MAX|ge MIN le MAX|ge MIN}
! Matches any routes with the subnet mask between 0-&gt;MAX, MIN-&gt;MAX or MIN-&gt;32 (le,le and ge,ge)
R(config)# ip prefix-list EXAMPLE 0.0.0.0/1 {le MAX|ge MIN le MAX|ge MIN}
! Matches any route to a Class A destination with a mask between 0-&gt;MAX, MIN-&gt;MAX or MIN-&gt;32
R(config)# ip prefix-list EXAMPLE 128.0.0.0/2 {le MAX|ge MIN le MAX|ge MIN}
! Matches any route to a Class B destination with a mask between 0-&gt;MAX, MIN-&gt;MAX or MIN-&gt;32
R(config)# ip prefix-list EXAMPLE 192.0.0.0/3 {le MAX|ge MIN le MAX|ge MIN}
! Matches any route to a Class C destination with a mask between 0-&gt;MAX, MIN-&gt;MAX or MIN-&gt;32
```

Additionally, the distribute list can be applied to all the BGP process:

```
R(config-router)# distribute-list [gateway] PREFIX-LIST {in|out} [INTERFACE]
! Use gateway to filter based on route source
```

### AS Path ACL

```
R(config-router)# neighbor NEIGH-ADDR filter-list AS-PATH-ACL {in|out}
```

To define an AS-PATH ACL, use:

```
R(config)# ip as-path access-list AS-PATH-ACL {permit|deny} REGEX
. = any character, including space
^ = beginning of string
$ = end of string
_ = ^$,() and space
\ = escape character
* = match zero or more occurances
+ = match one or more occurances
? = match zero or one occurances
! must be preceded by CTRL+V or ESC+Q to prevent ? character from being interpreted as HELP
```

A very useful command to test regular expressions is:

```
R# show ip bgp regexp REGEX
```

Some common REGEX values are:

```
.* - matches anything
^$ - matches empty string = Locally originated routes
^123_ - string that starts with 123 = Routes received from AS 123
_123$ - string that ends with 123 = Routes originated in AS 123
_123_ - string that contains 123 = Routes that passed through AS 123
^[0-9]+$ - string that only contains one number = Routes originated in a neighboring AS
\(([0-9]+)+\) - string that contains severa numbers between () = Routes that passed a confederation
```

Also check [this section](https://nyquist.eu/cisco-cli-tips-and-tricks/#21_Regular_Expressions) on regular expressions.

### Communities

Read about [COMMUNITY](/ipv4/ipv4-routing/bgp/bgp-attributes).

### Advertise Maps

You can do conditional advertisement of routes using an advertise map:

```
R(config-router)# neighbor NEIGH-ADDR advertise-map ADVERTISE-MAP {exist-map EXIST-MAP|non-exist-map NON-EXIST-MAP}
```

The routes matched in the ADVERTISE-MAP are advertised if the routes matched in the EXIST-MAP exist in the BGP table, or if the routes matched in the NON-EXIST-MAP do not exist in the BGP table.

## Summarization

### Auto Summary

Auto-summary only affects routes redistributed into BGP. By default, it is off. When on, it will summarize at the major network boundary

```
R(config-router)# auto-summary
```

### Aggregation

An aggregate router can be sent instead of it’s longer prefix length children.

```
R(config-router)# aggregate-address SUMMARY-ADDR SUMMARY-MASK OPTIONS
! available OPTIONS:
! as-set
! summary-only
! attribute-map ATTRIBUTE-MAP
! advertise-map ADVERTISE-MAP
! suppress-map SUPPRESS-MAP
```

* **as-set** = The default AS\_PATH of an aggregate is empty: $^. When using this keyword the router will add the AS\_SET to the AS\_PATH. AS\_SET is an unordered list of all ASes in the AS\_PATH of all children routes. It appears between {} in the AS\_PATH.
* **summary-only**= by default, both the aggregate and its children are advertised. Using this keyword only the summary is advertised, all other are suppressed
* **attribute-map ATTRIBUTE-MAP** = allows changing the attributes of the aggregate route
* **advertise-map ADVERTISE-MAP** = allows you to select only a subset of the routes that will be used to generate the aggregate. The routes that are not matched will be advertised separately and the aggregate will not inherit their attributes.
* **suppress-map SUPPRESS-MAP** = allows you to suppress a subset of the children routes. Routes that are matched by the route-map are suppressed

Routes that are suppressed can be unsuppressed per neighbor:

```
R(config-router)# neighbor NEIGH-ADDR unsuppress-map UNSUPPRESS-MAP
```

Routes that are matched in the UNSUPPRESS-MAP are advertised to the neighbor.

### Use a static route to NULL0

Another option is to add a static route for the summary to NULL0 and then add that network to BGP:

```
R(config)# ip route SUMMARY-ADDR SUMMARY-MASK NULL0
R(config)# router bgp AS-NUMBER
R(config-router)# network SUMMARY-ADDR [mask SUMMARY-MASK]
```

This will add the summary to the BGP advertised routes but will also advertise all children prefixes. Additional filtering of those children may be required.

## Default routes

### Originate network 0.0.0.0

To add default routes to the BGP table you can use the network command if a default route already exists in the routing table:

```
R(config-router)# network 0.0.0.0
```

### Redistribute routes to 0.0.0.0

Another option is to redistribute a route to 0.0.0.0 into BGP. It won’t be enough to make it work. You also have to let BGP know you want it to advertise the default route. The route could be known via an IGP or via static routes. For example:

```
R(config)# ip route 0.0.0.0 0.0.0.0 NULL0
! Have the default route in the routing table
R(config)# router bgp AS-NUMBER
R(config-router)# redistribute static [route-map ROUTE-MAP]
! Redistribute the default route into BGP
R(config-router)# default-information originate
! Let BGP knwo you want to advertise the default route
```

### Advertise a default route to a neighbor

A simpler approach is to use the following command:

```
R(config-router)# neighbor NEIGH-ADDR default-originate
```

This configuration will always advertise a default route to a neighbor, but you can also define a condition for advertisement using a route map:

```
R(config-router)# neighbor NEIGH-ADDR default-originate route-map CONDITION-MAP
```

## Redistribution

When redistributing from BGP to another protocol, only eBGP routes will be redistributed to prevent loops. Internal routes can be advertised using:

```
R(config-router)# bgp redistribute-internal
```

Take care when redistributing BGP in another protocol – the routes could be learned again from the IGP and usually this means that the IGP route will be preferred due to the lower AD. This may result in routing loops.


# BGP Attributes

## BGP Data Structures

### Neighbor Table

```
R# show ip bgp neighbors [NEIGH-ADDR]
R# show ip bgp summary
```

The address in the bgp summary table shows the IP used in the peering, not the Router ID.

### BGP Table

Lists all prefixes learned from all peers

```
R# show ip bgp [topology]
! Status codes:
! > = Best route - will be installed in the routing table
! r = RIB failure:
R# show ip bpg rib-failures
```

If no routes towards a destination show the “>” code, you should investigate why no route is considered valid. Check with:

```
R# show ip bgp PREFIX
```

## Route Attributes

### Well Known

Well Known attributes are those attributes that must be supported by every BGP implementation.

#### **Mandatory**

Mandatory attributes are those attributes that must be sent in each Update.

{% tabs %}
{% tab title="AS\_PATH" %}
Whenever a prefix is advertised to an eBGP peer, the AS is prepended to the AS\_PATH. AS\_PATH prepending is used to make the AS\_PATH longer than normal in order to make some routes more preferred than others.

**AS\_PATH prepending**

```
R(config-route-map)# set as-path prepend {AS-NUMBER AS-NUMBER... | last-as TIMES}
! Usually, the local AS is prepended multiple times
! When using last-as, the last AS, not the current AS is prepended several TIMES
```

**Ignoring AS\_PATH when searching best path**

The AS\_PATH is an important tiebreaker when selecting the best BGP route for a destination. However, it can be ignored with the hidden command:

```
R(config-router)# bgp bestpath as-path ignore
```

**AS\_PATH loop prevention**

AS\_PATH also acts as a Loop prevention mechanism. When a BGP peer receives a route that includes its own AS in the AS\_PATH, the update is dropped. An extreme way of filtering is to prepend the next AS to the AS\_PATH, forcing a neighbor to drop the route due to loop prevention.

However, there are situations when you would want to ignore this behavior and accept routes with your own AS-NUMBER in the AS\_PATH. (For example when you have the same AS number configured in 2 locations that connect over an ISP BGP cloud. Normally you would not accept routes that are originated in your own AS, but this time you should). To enable this feature, use:

```
R(config-router)# neighbor NEIGH-ADDR allowas-in TIMES
! accepts routes with the current AS in the AS-PATH. The AS may occur several TIMES in the AS_PATH
```

Another option si for the ISP to override the AS-NUMBER of the customer and replace it with its own AS in the AS-PATH. The command to do this is:

```
R(config-router)# neighbor NEIGH-ADDR as-override
! It will replace customer AS with the current router's AS
```

**Private AS**

When using Private ASes, you should strip them from the AS\_PATH when sending to an eBGP peer:

```
R(config-router)# neighbor NEIGH-ADDRESS remove-private-as
```

**Local AS**

Another way of modifying the AS\_PATH is via the **local-as** command, detailed [here](https://nyquist.eu/more-bgp/#5_Local_AS).

**AS\_PATH vs AS\_SET vs AS\_SEQUENCE vs AS\_CONFED\_SEQUENCE**

There are two types of AS\_PATH attributes:

* AS\_SET: an unordered list of AS numbers along a path. It is used when a BGP speaker creates an aggregate route. The AS\_SET will include all AS-es in the AS\_PATHs of the children routes in order to prevent loops. It will appear between {} in the AS\_PATH and will only count as 1 AS when calculating the shortest AS\_PATH. When the AS\_SET is included in the AS\_PATH, the ATOMIC\_AGGREGATE does not have to be included with the aggregate.
* AS\_SEQUENCE: an ordered list of AS numbers along a path
* When using confederations, an AS\_CONFED\_SEQUENCE appears in the AS\_PATH between (). They do not count in the Shortest AS\_PATH calculation.

{% endtab %}

{% tab title="ORIGIN" %}

Specifies the origin of the routing update:

* i = IGP – learned from an IGP. These routes are added to the BGP network using the “network” command
* e = EGP – learned from an EGP. You shouldn’t see this kind of routes in real life.
* ? = incomplete – usually as a result of redistribution. They are added to the BGP network using the “redistribute” command.

You can modify the origin of a route in a route-map, using:

```
R(config-route-map)# set origin {igp|egp AS-NUMBER|incomplete}
```

{% endtab %}

{% tab title="NEXT\_HOP" %}

When advertising to an eBGP neighbor, the NEXT\_HOP is set to the advertising router’s interface address. When advertising to an iBGP peer, the NEXT\_HOP is not modified, unless the next-hop-self command is used:

```
R(config-router)# neighbor NEIGH-ADDRESS next-hop-self
```

You can change the next-hop of sent or received routes in a route-map used in the outgoing or the incoming direction.

```
R(config-route-map)# set ip next-hop IP1 [IP2...]
! Sets the IP address of the incoming/outgoing routes to the first reachable IP.
R(config-route-map)# set ip next-hop peer-address
! For incoming route-maps, sets the next-hop to the neighbor-address, regardless of the received value
! For outgoing route-maps, it is simialr to next-hop-self, but only applies to
```

{% endtab %}
{% endtabs %}

#### **Discretionary**

Discretionary attributes are those attributes that are supported by every implementation of BGP (Well Known) but may not be sent in every Updated

{% tabs %}
{% tab title="LOCAL\_PREF" %}

Local Preference is used only in updates towards iBGP peers. If a BGP speaker receives multiple routes to the same destination, it compares the LOCAL\_PREF and the highest value wins. You can set the local preference in a route-map, using:

```
R(config-route-map)# set local-preference LOCAL_PREF
! default: 100
```

The default value of local-preference can be modified using:

```
R(config-router)# bgp default-local-preference LOCAL_PREF
! Default: 100
```

{% endtab %}

{% tab title="ATOMIC\_AGGREGATE" %}
Used to alert downstream routers that a loss of path information has occured due to summarization. When the ATOMIC\_AGGREGATE is used, the BGP speaker can also use the AGGREGATOR attribute.
{% endtab %}
{% endtabs %}

### Optional

Optional attributes are those attributes that may not be supported by every BGP implementation.

#### **Transitive**

Transitive attributes are those attributes that may not be supported by every BGP implementation (Optional) but must be passed to the neighbor in the Update messages.

{% tabs %}
{% tab title="AGGREGATOR" %}
It contains the AS number and the Router ID of the router that originated the Aggregate route

{% endtab %}

{% tab title="COMMUNITY" %}
A community represents a group of destinations which share one or more common properties. Or, in other words, communities are labels that are attached to certain routes. The format of the COMMUNITY attribute is in the form of one 32 bit number (1-4.294.967.295). There is also a “new-format”, where the 32 bit number is split into 2×16 bit sub-values.

```
R(config)# ip bgp-community new-format
```

The default cisco format for community is AS:NN, where AS is the AS number where the community was originated. The default IEEE format is NN:AS.

* 0 – 65535 (0x000000 – 0x0000FFFF) – reserved
* User Defined
* Well Known
  * INTERNET – all routes belong to this community by default
  * NO\_EXPORT (0xFFFFFF01) – these routes cannot be advertised to eBGP peers or to peers outside the confederation
  * NO\_ADVERTISE (0xFFFFFF02) – routes received with this value cannot be advertised at all, neither to eBGP nor iBGP peers
  * LOCAL\_AS (0xFFFFFF03) or NO\_EXPORT\_SUBCONF – routes received with this value cannot be advertised to eBGP peers. If a configuration is configured, the routes cannot be advertised to other SUB-AS-es within the confederation

To attach a community to a route and send it to a neighbor:

1. Use it in a route map:

   ```
   R(config-route-map)# set community COMMUNITY [additive]|COMMUNITY_LIST|none}
   ```
2. Apply the route-map on a neighbor:

   ```
   R(config-router)# neighbor NEIGH-ADDR route-map ROUTE-MAP out
   ```
3. Send community to the neighbor:

   ```
   R(config-router)# neighbor NEIGH-ADDR send-community
   ```

Once a community is set, it will overwrite all previous communities for that routes, unless the **additive** keyword is used.

A community list is used to match one or more Communities. To create a community-list, use:

* Standard

  ```
  R(config)# ip community-list COMMUNITY-LIST {permit|deny} COMMUNITY
  ! COMMUNITY-LIST can be a string or a number (1 - 99)
  ```
* Expanded

  ```
  R(config)# ip community-list COMMUNITY-LIST {permit|deny} REGEX
  ! COMMUNITY-LIST can be a string or a number (100 - 500)
  ```

You can match by a single community or by a Community list:

```
R(config-route-map)# match community COMMUNITY_LIST [exact-match]
! by default the COMMUNITY_LIST is matched using OR
! exact-match - use AND when matching
R(config-route-ma)# match community COMMUNITY
```

{% endtab %}
{% endtabs %}

#### **Non-Transitive**

Non-Transitive attributes are those attributes that may not be supported by every BGP implementation (Optional) are are not be passed to the neighbor in the Update messages.

{% tabs %}
{% tab title="MED (MULTI\_EXIT\_DISC)" %}

Allows an AS to inform another AS of the preferred ingress point. The lowest MED is preferred. The MED attribute is not passed beyond the receiving AS, so it is used to influence traffic between 2 directly connected AS-es. To influence route preference beyond the neighboring AS, use AS\_PATH prepend. If an ISP must forward traffic to a client AS, it will usually try to find the shortest way to send the traffic to the client AS. This behavior is cold “Hot Potato Routing”. However, the client might prefer this traffic to enter it’s network on a router that it’s closer to the actual destination. In this way, known as “Cold Potato Routing” the traffic will travel more inside the ISP and less inside the Client’s network. Routing by MED is one way of implementing “Cold Potato Routing”. By default, MED is compared only for routes that come from the same AS. Use the following command to compare MED regardless of the AS they were received:

```
R(config-router)# bgp always-compare-med
```

By default, routes are compared in order, the newest to the 2nd newest, the best of these to the third newest and so on to the oldest route. When comparing MED, this order can be changed using:

```
R(config-router)# bgp deterministic-med
```

When this command is enabled, the routes are grouped by the last AS and a best route is chosen for each group. if the “bgp always-compare-med” command is also used, the best routes for each group are compared between them and the lowest MED wins. Else, the MED will not be used because the routes came from different ASes and the route decision processes will go to the next step. See [this article](https://www.cisco.com/en/US/tech/tk365/technologies_tech_note09186a0080094925.shtml) for an example. By default, a Cisco router considers a route that doesn’t have the MED attribute set as heaving the value 0, which is the most preferred MED. You can change this behavior, using:

```
R(config)# bgp bestpath med missing-as-worst
```

This will make the router consider a route with a missing MED as the worst MED (4.294.967.295) You can enable MED comparison only for routes originated in the same confederation, using:

```
R(config-router)# bgp bestpath med-confed
```

When redistributing into BGP, the router will use the value set for the IGP metric as the value of the MED attribute of the route. If no value is set in the redistribution command, then the value of default-metric is used:

```
R(config-router)# default-metric MED
! Default: 0
```

{% endtab %}

{% tab title="ORIGINATOR\_ID" %}
Contains the ROUTER\_ID of the originator of the route in the local AS (the last eBGP peer). A router should ignore routes received with its ROUTER\_ID as the ORIGINATOR\_ID. This attribute is set by the Route Reflector when reflecting routes to its clients.
{% endtab %}

{% tab title="CLUSTER\_LIST" %}
A route reflector must prepend the local CLUSTER\_ID to the CLUSTER\_LIST. CLUSTER\_ID is the same as the ROUTER\_ID unless specifically changed:

```
R(config-router)# bgp cluster-id CLUSTER-ID
```

When a router receives a route with its Cluster ID in the Cluster List, the route is discarded.
{% endtab %}
{% endtabs %}

### Weight

It is a proprietary Cisco attribute that applies only to the current router and is not advertised to other peers. The higher the weight, the more preferable the route. By default, routes learned from peers have a weight of 0, while locally generated routes have a weight of 65535. To set the Weight of incoming routes, use:

```
! For a neighbor
R(config-router)# neighbor NEIGH-ADDR weight WEIGHT
! In a route-map
R(config-route-map)# set weight WEIGHT
```

## Choosing the best route

### Valid routes

In order to enter the route selection process, a route received via BGP must be considered valid. To do so, first it should be verified if the route appears in the BGP table:

```
R# show ip bgp
! * - Valid Route
! r - RIB Failure
! s - Suppressed
! d - Damped
```

Most often then not, the reasons why a route is not considered valid, are:

* If BGP Synchronization is enabled, the route is not found in the IGP table. When using OSPF, the Router ID of OSPF and BGP must match
* The router ignores the prefix because the current AS-NUMBER is already in the AS-PATH of the received route – eBGP only
* The route is dampened
* The next-hop is unreachable

For each prefix, the route that is chosen as the best route will be marked with the “>” character.

### Route Selection Process

1. **W** – Highest **W**eight – Cisco only
2. **L** – Highest **L**OCAL\_PREF
3. **OL** – Prefer a route that was **O**riginated **L**ocally. Routes originated with the network command are preferred to routes originated with the redistribute command, which are preferred to routes originated with the aggregate command, which are preferred to routes learned from other peers.
4. **A** – Shortest **A**S\_PATH
5. **O** – Lowest **O**rigin Code = i < e < ?
6. **M** – Lowest **M**ED if the routes go to the same AS
7. **E** – Prefer **e**BGP routes first, then iBGP routes
8. **N** – Shortest path to the **N**EXT\_HOP – lowest IGP (AD/metric) to the NEXT\_HOP
9. **M** – If **M**ultipath is enabled, install multiple routes in the routing table, if they are from the same neighboring AS:

   ```
   R(config-router)# maximum-paths [ibgp] PATHS
   ! PATHS: 1 - 6. Default: 1
   ```

   Although multiple paths are added to the routing table, only one route will be considered best path and will be advertised to other neighbors. If Multipath is disabled, or no best route has been selected, then continue to the next step
10. **O** – Oldest routes – only for eBGP routes. There are situations when this step is skipped, like when using the command:

    ```
    R(config-router)# bgp bestpath compare-routerid
    ```

    See [this article](https://www.cisco.com/en/US/tech/tk365/technologies_tech_note09186a0080094431.shtml).
11. **R** – Prefer the route originated by the router with the lowest **R**outer ID (or Originator ID in Route Reflector environments).
12. **C** – For routes originated by the same Router ID, choose the one with the shortest **C**LUSTER\_LIST
13. **N** – Prefer the route that came from the **N**eighbor with the lowest address

Remember it as **With L.OL.A. O.M.E.N. M.O.R. C.N.**. For detailed information, read [this article](https://www.cisco.com/en/US/tech/tk365/technologies_tech_note09186a0080094431.shtml)


# More BGP

## Route Dampening

It is used to stop unstable routes from being forwarded throughout the network. When a route flaps, a penalty is assigned to the route (Default: 1000 per flap). A timer called Half-Life is used to reduce the penalty value to half (Default: 15 min). If the penalty value exceeds the suppress limit, the route is no longer advertised (Default: 2000). The route continues to be suppressed until the Half-Life reduces the penalty below the reuse limit (Default: 750). A route cannot be suppressed more than the Maximum Suppress Time (60 min or 5 \* Half-Life).\
Dampening can be enabled globally:

```
R(config-router)# bgp dampening [PARAMS]
```

or only for some routes matched in a route-map:\
R(config-router)# set dampening PARAMS\
! Must set PARAMS. It won’t use default settings\
When setting dampening with a route-map, define dampening parameters in the route-map.\
You can verify dampening with:

```
R# show ip bgp dampening parametrs
```

Dampened prefixes can be manually cleared, using:

```
R# clear ip bgp dampening PREFIX NETMASK
```

However, this will not clear the penalty value so if a new flap occurs, the route will probably be immediately dampened. To clear the penalty value, use:

```
R# clear ip bgp NEIGH-ADDR flap-statistics
```

## Backdoor networks

By default, an external BGP learned route is preferred due to the lower AD (20) to an IGP learned route. If you would like to prefer the route learned via an IGP, use the Backdoor command on the network:

```
R(config-router)# network NETWORK-ADDR backdoor
```

This command will modify the AD of the BGP learned route to 200, making it less preferred over other IGP routes.

## Fast Fallover

### Fast External Fallover

For eBGP peers that are directly connected, the router will bring down the neighbor relationship if the interface status goes down. In case of flapping interfaces, you may want to keep the neighbor relationship up until the dead timer expires.

```
R(config-router)# [no] bgp fast-external-fallover
! Default: on
```

The command can be issued per interface, with:

```
R(config-router)# ip bgp fast-external-fallover {permit|deny}
! permit - enables fast-external-fallover
! deny - disables fast-external-fallover
```

### Internal Fallover

For iBGP peers you can also enable fast fall-over with the command:

```
R(config-router)# neighbor NEIGH-ADDR fall-over [route-map ROUTE-MAP]
! ROUTE-MAP can be used for conditional deactivation of a session
```

However, this is usually not a desired behavior.

## ORF – Outbound Route Filtering

Sending a lot of updates which are filtered inbound by a neighbor is unnecessary but there was no way for a router to know how its neighbors would handle the routes. With the introduction of the ORF feature, 2 ORF-capable router can exchange information regarding their inbound filters, so that the sending router can filter them in the outbound direction.\
To enable this feature, 2 neighbor routers must be configured with:

```
R(config)# neighbor NEIGH-ADDR capability orf prefix-list {both|send|receive}
! both     Capability to SEND and RECEIVE the ORF to/from this neighbor
! receive  Capability to RECEIVE the ORF from this neighbor
! send     Capability to SEND the ORF to this neighbor
```

After the command is enabled on both routers, they will exchange information and will filter outbound updates with the same prefix list as the filter set on the other peer in the inbound direction, therefor optimizing the bandwidth usage.

## Local AS

For temporary situation when a company migrates from one AS number to another, it is necessary to be able to change the AS number used on a per-neighbor basis.\
This can be done using the command:

```
R(config-router)# neighbor NEIGH-ADDR local-as NEW-AS [no-prepend [replace-as [dual-as]]]
```

Now, the router will make connections to this neighbor as if it is running BGP in the NEW-AS.

When the command is entered only with the local-as keyword, then the router will do the following:

* The AS\_PATH of the routes received from this neighbor will be prepended with the NEW\_AS. When they are sent out to other eBGP peers, they will also be prepended with the OLD\_AS
* The AS\_PATH of the routes sent to this neighbor will be prepended with both the NEW\_AS and the OLD\_AS, with the NEW\_AS appearing first in the AS\_PATH

When the command is entered with the “no-prepend” keyword, the router will do the following:

* The AS\_PATH of the routes received from this neighbor will not be prepended with the NEW\_AS. When they are sent out to other eBGP peers, they will only be prepended with the OLD\_AS
* The AS\_PATH of the routes sent to this neighbor will be prepended with both the NEW\_AS and the OLD\_AS, with the NEW\_AS appearing first in the AS\_PATH

When the command is entered with the “no-prepend replace-as” keywords, the router will do the following:

* The AS\_PATH of the routes received from this neighbor will not be prepended with the NEW\_AS. When they are sent out to other eBGP peers, they will only be prepended with the OLD\_AS
* The AS\_PATH of the routes sent to this neighbor will be prepended only with the NEW\_AS, skipping the OLD\_AS

When using also the “dual-as” keyword, the router will accept peering with this neighbor on both AS-NUMBERS, making it easy to migrate from one AS-NUMBER to another.

## Maximum prefixes

The internet BGP tabel size is huge and if you receive such a large number of routes from a neighbor, you’re router might not have the amount of memory to manage it. This is why you can configure a maximum number of prefixes that you are willing to receive from your neighbor. When the Maximum number of prefixes is reached, the session is shut down.

```
R(config-router)# neighbor NEIGH-ADDR maximum-prefix MAX [TH] [restart TIMER] [warning-only]
! MAX = maximum number of routes
! TH = When the number of received routes reaches TH% of MAX, the router generates a warning. Default: 75
! restart TIMER = The session will be reestablished after the TIMER expires. If not configured, it will remain shutdown.
! warning-only = doesn't shutdown, but issues a warnign (syslog)
```


# Route Redistribution

## Route Redistribution

You can redistribute routes from one routing process to another using the redistribute command inside the destination routing process:

```
R(config-router)#redistribute SRC [PROC|AS] [metric METRIC] [route-map ROUTE-MAP] [OSPF-OPTIONS]
! SRC: rip, eigrp, ospf, isis, bgp, static, connected, odr
! PROC|AS: used for OSPF and EIGRP
! METRIC: seed metric into the DESTINATION process
! ROUTE-MAP: used for conditional redistribution
! OSPF-OPTIONS: options specific to OSPF
```

When a routing protocol starts, it automatically redistributes connected routes that are matched by the network command. This also happens for static routes that point to an interface if they are matched by a network command. These routes are considered to be directly connected and will be redistributed without the **static** keyword.

### Redistributing into OSPF

When redistributing into OSPF, automatic classless summarization takes place, unless the **subnets** keyword is specified:

```
R(config)# router ospf PROC
R(config-router)# redistribute [SRC [PROC|AS]] subnets
```

### Redistributing from OSPF

When redistributing OSPF routes, you can select the types of routes that are redistributed.

```
R(config-router)# redistribute ospf PROC match {internal|external {1|2}|nssa-external {1|2}}
! using just external means both 1 and 2
```

When redistributing OSPF routes into BGP, the default behavior is to redistribute only OSPF intra-area and inter-area routes. You must use the **external** keyword to enable the redistribution of OSPF external routes.

### Redistributing BGP into IGP

When redistributing BGP, only eBGP routes are redistributed by default. You can modify this behavior if you use the command:

```
R(config-router)# bgp redistribute-internal
```

## Seed metric

By default only when redistributing from one EIGRP process to another or from one OSPF process to another, the metric values can be used from the first to the second.

### RIP

When redistributing from any other protocol to RIP, the **metric transparent** keyword can make RIP use the numerical value of the metric from the other protocol. If this value is greater than 15 (usually), the route will be considered unreachable in RIP.

```
R(config)# router rip
R(config-router)# redistribute [SRC [PROC|AS]] metric {transparent|VALUE}
```

The router that performs redistribution into RIPt will send updates to other RIP routers using the seed metric. When those routers will send updates to the next hop, they will add one more hop to the metric.

### OSPF

When redistributing from any other protocol to OSPF, the default metric value is 20 with a metric type of E2.

```
R(config)# router ospf
R(config-router)# redistribute [SRC [PROC|AS]] metric VALUE [metric-type {type1|type2}]
```

### EIGRP

When redistributing from any other protocol to EIGRP, the metric cannot be auto-converted. You must use the EIGRP **metric** keyword with the redistribute command.

```
R(config)# router eigrp
R(config-router)# redistribute [SRC [PROC|AS]] metric BW DEL REL LOAD MTU
```

### Default metric

You can alse set a default metric regardless of the route source:

```
R(config-router)# default-metric {VALUE|BW DEL REL LOAD MTU}
! BW DEL REL LOAD MTU: used for EIGRP
! VALUE: used for other protocols
```

The metric used with the redistribute command overrides this default value.

## Conditional Redistribution

Conditional redistribution can be achieved by using a route-map. Only routes that are allowed in the route-map are redistributed into the protocol. Also, the set commands in the route-map are applied to the redistributed routes.\
A route map is defined as a series of entries identified by a sequence number (SEQ). Routes that are matched by a permitting entry are redistributed, while routes that are matched by a denying entry are not redistributed.

### Match rules

Each route is passed through each entry until a match is found, if no match is found, the default is to not redistribute the route. You can change this by using an entry with no match commands that would match any route.

```
R(config)# route-map REDIST-MAP {permit|deny} [SEQ]
! default: permit 10
```

In each entry a set of match rules define the routes that are matched by the entry. If multiple match rules are defined, a packet must pass all of them (logical AND). Some match rules can have multiple entries (for example multiple ACLs in one line). In this case, a route is considered to be matched if it is matched by at least one entry in the match rule (logical OR). Also, by entering the same match type with different conditions (e.g. multiple **match ip address ACL**) they will be converted to one match entry with multiple entries – therefore the logical OR will apply.

#### **General match rule**

```
R(config-route-map)# match interface INTERFACE
! Matches routes that have the next hop on the defined INTERFACE
R(config-route-map)# match {ip|ipv6} address ACL
! Matches routes that point to a destination matched by the ACL
R(config-route-map)# match {ip|ipv6} prefix-list PREFIX-LIST
! Matches routes that point to a destination matched by the PREFIX-LIST
! A PREFIX-LIST can also match the network mask
R(config-route-map)# match {ip|ipv6} next-hop ACL
! Matches routes that have a next-hop matched by the ACL
R(config-route-map)# match {ip|ipv6} route-source {ACL|prefix-list PREFIX}
! Matches routes advertised by an IP matched in the ACL/prefix-list
R(config-route-map)# match tag TAG
! Matched routes tagged with TAG
R(config-route-map)# match metric VALUE [+- D]
! Matches routes with a specific VALUE or in the range VALUE-D VALUE+D
```

#### **IGP specific**

We saw before that we can use an extended ACL to match a route using:

```
R(config-route-map)# match {ip|ipv6} address ACL
! Matches routes that point to a destination matched by the ACL
```

When dealing with IGP routes, the extended ACL will be interpreted as follows: The SOURCE in the extended ACL will match the update source, while the DESTINATION will match the route destination.\
E.g:

```
R(config)# access-list 101 permit host SRC host DST
! will match routes advertised by SRC for the destination DST
```

Of course, any binary math can still be used.\
Other rules:

```
R(config-route-map)# match route-type {external|internal|nssa-external}
! OSPF only
R(config-route-map)# match route-type external
! EIGRP only
```

#### **BGP specific**

We saw before that we can use an extended ACL to match a route using:

```
R(config-route-map)# match {ip|ipv6} address ACL
! Matches routes that point to a destination matched by the ACL
```

When dealing with BGP routes, the extended ACL will be interpreted as follows: The SRC in the extended ACL will match the destination address, while the DST will match the network mask.\
E.g:

```
R(config)# access-list 101 permit host 10.0.0.0 host 255.255.255.0
! will match the route 10.0.0.0/24
```

Of course, any binary math can still be used.

```
R(config-route-map)# match route-type local
! Locally originated BGP routes
R(config-route-map)# match as-path PATH-LIST
R(config-route-map)# match community COMM [exact-match]
R(config-route-map)# match local-preference LP
```

### Set rules

Set commands:

```
! Set metric
R(config-route-map)# set metric {VALUE|BW DEL REL LOAD MTU}
! Set OSPF metric-type
R(config-route-map)# set metric-type {type-1|type-2}
! Set automatic tags
R(config-route-map)# set automatic-tag
! Set manual tag
R(config-route-map)# set tag TAG
! BGP Specific:
R(config-route-map)# set community COMMUNITY
R(config-route-map)# set local-preference LP
R(config-route-map)# set weight WEIGHT
R(config-route-map)# set origin {igp|egp AS|incomplete}
R(config-route-map)# set as-path [prepend] AS-LIST
```


# Policy based Routing

## In Theory

Policy Routing is a feature on Cisco routers that offers greater control over paths in a network.\
Normal routing is done based on the routing table. The router looks up the destination address of a packet and finds out the next hop or the outgoing interface.\
With Policy Routing you can override whatever the routing table says and just route packets however you like.\
Isn’t this what static routing does? To some degree, yes. You can use static routes to manipulate the outcome of a route lookup, but static routes integrate in the normal routing process. The router will still base it’s routing decision on the packet’s destination address.\
The true power of policy routing is that it can take routing decisions on other parameters like incoming interface, source address or even application type (http, ftp).

Configuring Polcy Routing is a 2 step process.

### Define the route-map

```
R(config)# route-map ROUTE-MAP [permit|deny] [SEQ]
!default: permit 10
```

A route-map can have multiple statements, each statement being identified by a **SEQ**. If you don’t set a number, the default is 10. One statement can have multiple match commands and multiple set commands.

A packet is passed sequentially through the route-map statements. At each statement, the packet is considered matched if it is matched by all the match commands in that statement. If at least one of the match statements is not matched, then the packet is passed to the next statement. If all match statements are matched, then all set commands are applied.

If a matched route-map statement is set to permit, then the packet is policy-routed. If a matched route-map statement is set to deny, then it is routed according to the routing table. In either case, once a matching statement is found, no other statements are checked. If no statements are matched the packet is also routed according to the routing table.

Inside the route-map configuration mode we have to define a matching criteria using the **match** command. There are a lot of options, but most of them have to do with manipulating routing updates. For the purpose of Policy Routing we can use the following commands. In one statement, there can be multiple match commands, and a packet will be considered to match if it matches all of them (AND logic). One match command can refer to multiple objects (ex: multiple ACLs). A packet will be considered to match that entry if it matches at least one of the objects (OR logic).

```
R(config-route-map)# match interface INTERFACE
!Matches on incoming interface:
R(config-route-map)# match ip address ACL
!Matches on destination IP:
R(config-route-map)# match ipv6 address ACL
!Matches on IPv6 destination IP:
R(config-route-map)# match length MIN-LEN
!Matchs on minimum packet length:
```

When a packet is matched, all the **set** commands that are configure in the statement are executed:

```
R(config-route-map)# set [default] interface INTERFACE
! Sets the outgoing interface
R(config-route-map)# set ip [default] next-hop NEXT-HOP-IP [recursive NEXT-HOP-IP2]
! Sets the next-hop address (IPv4)
R(config-route-map)# set ipv6 [default] next-hop NEXT-HOP-IP
! Sets the next-hop address (IPv6)
```

When using the **default** keyword we tell the router to first try to find a match using the routing table and then use our defined policy, basically overriding only the routing table’s default route.

As you can see, we can define an outgoing interface or a next-hop address. Just like with static routes, you must understand the differences between using the next-hop or the outgoing interface. If there are multiple set commands, the order in which they are resolved is:

1. set ip next-hop
2. set ip next-hop recursive
3. set interface
4. set ip default next-hop
5. set ip default interface

For IPv4 we have one more option – using EOT – Enhanced Object Tracking to see if the next hop is up or not:

```
R(config-route-map)#set ip next-hop verify-availability NEXT-HOP-IP SEQ track OBJ
```

**SEQ** is a sequence number that is used to identify multiple options for the next hop, end **OBJ** is the object used to track availability. As you can imagine, this next hop will be used only if the object used for tracking is returning a positive result.

Optionally, you can also set a description for the route-map entry

```
R(config-route-map)# description STRING
```

#### **Use route maps for traffic filtering**

Simply match incoming traffic and set the outgoing interface to NULL0:

```
R(config-route-map)#set interface NULL0
```

### Apply the Route Map

You probably realized that the policy must be set on the interface where we expect to receive the traffic that will be matched. Setting the route-map on the outgoing interface makes no sense because the packets have already went through the routing process so they will just be sent out.

```
R(config-if)# ip policy route-map ROUTE-MAP
```

But what if I want to policy route traffic originated from the router. It has no incoming interface, so we have to use:

```
R(config)# ip local policy route-map ROUTE-MAP
```

#### **Force local traffic to be checked by ACLs**

By default, traffic originated locally is not checked by ACLs set on interfaces. You can use route-maps to set the next-hop of locally originated traffic to a loopback interface, and force this way the traffic from being checked by the ACLs. The loopback interface is still local but the fact that is switching from one interface to another will force the router to apply the ACLs on this traffic.

## Working examples

### Example 1

Here’s an example:

[![](https://nyquist.eu/wp-content/uploads/2012/03/PBR.png)](https://nyquist.eu/wp-content/uploads/2012/03/PBR.png)

PBR Example Topology

In our topology, we have 5 routers, each configured with an ip address on the Lo0 interface. On R1 we will simulate 2 hosts by using Lo1, as well. We will run EIGRP for route exchange. Here’s the starting config:\
On R1:

```
interface Loopback0
 ip address 1.1.1.1 255.255.255.255
!
interface Loopback1
 ip address 11.11.11.11 255.255.255.255
!
interface FastEthernet0/0
 ip address 12.0.0.1 255.255.255.0
 duplex auto
 speed auto
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
```

On R2:

```
interface Loopback0
 ip address 2.2.2.2 255.255.255.255
!
interface FastEthernet0/0
 ip address 12.0.0.2 255.255.255.0
 duplex auto
 speed auto
!
interface FastEthernet0/1
 ip address 24.0.0.2 255.255.255.0
 duplex auto
 speed auto
!
interface Serial1/0
 ip address 23.0.0.2 255.255.255.0
 serial restart-delay 0
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
```

On R3:

```
interface Loopback0
 ip address 3.3.3.3 255.255.255.255
!
interface FastEthernet0/1
 ip address 35.0.0.3 255.255.255.0
 duplex auto
 speed auto
!
interface Serial1/0
 ip address 23.0.0.3 255.255.255.0
 serial restart-delay 0
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
```

On R4:

```
interface Loopback0
 ip address 4.4.4.4 255.255.255.255
!
interface FastEthernet0/0
 ip address 45.0.0.4 255.255.255.0
 duplex auto
 speed auto
!
interface FastEthernet0/1
 ip address 24.0.0.4 255.255.255.0
 duplex auto
 speed auto
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
```

On R5:

```
interface Loopback0
 ip address 5.5.5.5 255.255.255.255
!
interface FastEthernet0/0
 ip address 45.0.0.5 255.255.255.0
 duplex auto
 speed auto
!
interface FastEthernet0/1
 ip address 35.0.0.5 255.255.255.0
 duplex auto
 speed auto
!
router eigrp 100
 network 0.0.0.0
 no auto-summary
```

Remember that by default, physical router interfaces are disabled, so you should enable them using

```
R(config-if)# no shutdown
```

On R1 we can test connectivity to R5 and see how the network converged:

```
R1#traceroute 5.5.5.5 source 1.1.1.1 numeric

Type escape sequence to abort.
Tracing the route to 5.5.5.5

  1 12.0.0.2 8 msec 24 msec 16 msec
  2 24.0.0.4 28 msec 36 msec 48 msec
  3 45.0.0.5 44 msec *  88 msec

R1#traceroute 5.5.5.5 source 11.11.11.11 numeric

Type escape sequence to abort.
Tracing the route to 5.5.5.5

  1 12.0.0.2 12 msec 24 msec 16 msec
  2 24.0.0.4 36 msec 36 msec 40 msec
  3 45.0.0.5 72 msec *  56 msec
```

Use **numeric** when starting a traceroute to disable the DNS lookups and get a faster response. As you can see, the route goes through R2, then R4 and then R5. The reason for this is because on R2, the interface towards R4 is Fast Ethernet, while the link towards R3 is over a HDLC serial link. Since the default EIGRP metric involves bandwidth and delay the route towards R4 will be considered the best. If we had RIP running, it would have installed 2 equal routes towards R5, as the hop count is the same.

```
R2#sh int fa0/1 | i BW
  MTU 1500 bytes, BW 10000 Kbit, DLY 1000 usec,
R2#sh int s1/0 | i BW
  MTU 1500 bytes, BW 1544 Kbit, DLY 20000 usec,
```

Now let’s set a route map on R2 so that traffic from 11.11.11.11 would go through R3:

```
!Define an ACL to match the traffic:
R2(config)# access-list 99 permit 11.11.11.11
!Define the route map
R2(config)# route-map ROUTE_MAP11 permit 10
R2(config-route-map)# match ip address 99
R2(config-route-map)# set ip next-hop 23.0.0.3
R2(config-route-map)# exit
!Apply the policy on the incoming interface:
R2(config)# interface Fa0/0
R2(config-if)# ip policy route-map ROUTE_MAP11
```

To verify the configuration, we will traceroute from R1 using each Loopbacks as the source:

```
R1#traceroute 5.5.5.5 source 1.1.1.1 numeric

Type escape sequence to abort.
Tracing the route to 5.5.5.5

  1 12.0.0.2 12 msec 16 msec 20 msec
  2 24.0.0.4 40 msec 40 msec 44 msec
  3 45.0.0.5 52 msec *  68 msec

R1#traceroute 5.5.5.5 source 11.11.11.11 numeric

Type escape sequence to abort.
Tracing the route to 5.5.5.5

  1 12.0.0.2 12 msec 20 msec 20 msec
  2 23.0.0.3 48 msec 64 msec 72 msec
  3 35.0.0.5 60 msec *  72 msec
```

As you can see, when the source is 1.1.1.1, the path is R1-R2-R4-R5and when the source is 11.11.11.11, the path is R1-R2-R3-R5\
To verify on R2, we can start a debug process to see how the traffic was routed.

```
R2#debug ip policy
Policy routing debugging is on
R2#
*Mar  1 02:44:30.919: IP: s=1.1.1.1 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy rejected(no match) - normal forwarding
*Mar  1 02:44:30.955: IP: s=1.1.1.1 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy rejected(no match) - normal forwarding
*Mar  1 02:44:30.999: IP: s=1.1.1.1 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy rejected(no match) - normal forwarding
*Mar  1 02:44:31.043: IP: s=1.1.1.1 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy rejected(no match) - normal forwarding
*Mar  1 02:44:31.095: IP: s=1.1.1.1 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy rejected(no match) - normal forwarding
*Mar  1 02:44:34.095: IP: s=1.1.1.1 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy rejected(no match) - normal forwarding
*Mar  1 02:44:44.463: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy match
*Mar  1 02:44:44.463: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, g=23.0.0.3, len 28, FIB policy routed
*Mar  1 02:44:44.523: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy match
*Mar  1 02:44:44.523: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, g=23.0.0.3, len 28, FIB policy routed
*Mar  1 02:44:44.583: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy match
*Mar  1 02:44:44.587: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, g=23.0.0.3, len 28, FIB policy routed
*Mar  1 02:44:44.647: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy match
*Mar  1 02:44:44.647: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, g=23.0.0.3, len 28, FIB policy routed
*Mar  1 02:44:44.707: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy match
*Mar  1 02:44:44.707: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, g=23.0.0.3, len 28, FIB policy routed
*Mar  1 02:44:47.679: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, len 28, FIB policy match
*Mar  1 02:44:47.679: IP: s=11.11.11.11 (FastEthernet0/0), d=5.5.5.5, g=23.0.0.3, len 28, FIB policy routed
```

Or we can use this command to see the policies defined and how many packets were matched:

```
R2#sh route-map
route-map ROUTE_MAP11, permit, sequence 10
  Match clauses:
    ip address (access-lists): 99
  Set clauses:
    ip next-hop 23.0.0.3
  Policy routing matches: 6 packets, 360 bytes
```

Notice there are 6 packets that matched, because although we sent 12 packets (check the debug output), only 6 of them matched the criteria – those sourced by 11.11.11.11. Why do we have 6 packets for just one traceroute command?\
Traceroute uses the TTL field to discover the path to a destination. It first sends packets with TTL=1, then increases the TTL until it reaches the destination. Each router on the path should decrease the TTL when routing the packet and should drop it when the TTL reaches 0. Actually, if the packet is not for itself, it would be pointless to route it, then decrease TTL and drop it, So, incoming packet with TTL 1 that should be routed are dropped. When dropping, it also should send an ICMP reply with the Type 11 (time-exceeded) and Code 0 (time to live exceeded in transit) back to the source, thus helping traceroute to map each router along the path.\
Cisco routers send by default 3 UDP packets for each TTL value starting from 1 upwards. This will help find multiple routes to a destination and also compute some average values for the Round Trip Time to each router along the path.\
So R1 first sends 3 packets with TTL=1. They will be droped by R2, so they won’t even get to the routing part – Matches so far: 0\
R1 then sends another 3 packets with TTL=2. They will pass through R2 before being dropped by R3 – Matches so far: 3\
R1 then sends another 3 packets with TTL=3. They will pass through R2 and R3 before R4 responds with an ICMP Type 3 (Destination unreachable) Code 3 (Port unreachable) for each one, signaling that the packets reached the destination, but there is no one waiting this kind of traffic. Matches so far: 6\
So there it is, the 6 matches that we saw, made sense.

### Example 2

Let’s say I want to have an out-of-band management connection from R2 to R5, when connecting via telnet. Let’s enable telnet on R5 first:

```
R5(config)# line vty 0 4
R5(config-line)# password cisco
R5(config-line)# login
```

We should be able to telnet from R2 to R5, but we are going through R4. I am using 2.2.2.2 as the source address of my telnet traffic.

```
R2#telnet 5.5.5.5 /source-interface lo0
```

Let’s set this management traffic to be out-of-band, through R3:

```
! Create the ACL:
R2(config)# access-list 199 permit tcp any any eq telnet
!Create the route-map
R2(config)# route-map ROUTE_MAP_TELNET permit 10
R2(config-route-map)# match ip address 199
R2(config-route-map)# set ip next-hop 23.0.0.3
! Apply the route-map
R2(config)# ip local policy route-map ROUTE_MAP_TELNET
```

In order to verify, let’s create an ACL on R5 just for accounting purpose:

```
! First allow telnet packets for accounting
R5(config)# access-list 100 permit tcp any any eq telnet log-input
! Then allow anything else
R5(config)# access-list 100 permit ip any any
! Apply it on Fa0/1 interface
R5(config)# interface Fa0/1
R5(config-if)# ip access-group 100 in
```

Back on R2, we should be able to telnet to R5, but remember to disable debugging first, otherwise we would see a lot of messages like this:

```
*Mar  1 06:11:53.994: IP: route map ROUTE_MAP_TELNET, item 10, permit
*Mar  1 06:11:53.994: IP: s=2.2.2.2 (local), d=5.5.5.5 (Serial1/0), len 40, policy routed
```

It’s just the **debug ip policy** that was previously started, telling us that there are locally generated packets that match the route-map.\
Over to R5 to verify:

```
R5#sh access-lists 100
Extended IP access list 100
    10 permit tcp any any eq telnet log-input (57 matches)
    20 permit ip any any (265 matches)
```

Yes! The telnet traffic originated by R2 arrived on Fa0/1 via R3.

### The end?

We were able to send traffic from R2 to R5, via R3, using the route-map. But is it over? Well, no, because this is just one way. If we look at R5’s routing table, it will send return traffic back to R2’s Lo0 via R4:

```
R5#sh ip route 2.2.2.2
Routing entry for 2.2.2.2/32
Known via "eigrp 100", distance 90, metric 435200, type internal
Redistributing via eigrp 100
Last update from 45.0.0.4 on FastEthernet0/0, 01:06:22 ago
Routing Descriptor Blocks:
* 45.0.0.4, from 45.0.0.4, 01:06:22 ago, via FastEthernet0/0
Route metric is 435200, traffic share count is 1
Total delay is 7000 microseconds, minimum bandwidth is 10000 Kbit
Reliability 255/255, minimum MTU 1500 bytes
Loading 1/255, Hops 2
```

The solution is to enable PBR in the other direction also. You should know how to do it by now.


# PfR 101 – Perfromance Routing

## PfR Technology

PfR stands for Performance Routing, but the feature was first called OER (Optimized Edge Routing). This is why most commands still start with the **oer** keyword. The idea behind PfR is to have a controlling entity (Master Controller) that takes over routing decisions for one or more Border Routers, in order to deliver traffic according to a specified policy. The MC will monitor network statistics (NetFlow, IP SLA) and will chose the best exit point for each type of traffic. PfR is used to enhance the operation of traditional routing protocols by using real time network performance information.

## PfR Components

Each PfR implementation requires the following minimum components:

* 1x Border Router
* 1x Master Controller
* 1x Internal Interface
* 2x External Interfaces

The MC and the BR can be configured on the same device. They both need to have a matching key-chain defined, that will secure communication between them, even if running on the same device:

```
R(config)# key chain PFR-KEY
R(config-key)# key KEY-ID
R(config-key-string)# key-string PASSWORD
```

### Border Router

To define a router as OER BR, use:

```
R(config)# oer border
```

Inside the **oer-br** configuration mode, you can define the interface used for connection to the MC and the IP address of the MC:

```
R(config-oer-br)# local INTERFACE 
! usually Loopback
R(config-oer-br)# master MASTER-IP key-chain PFR-KEY
```

Normally, OER uses TCP port 3949 on both the MC and BR. This can be changed with the following command, but it has to match on both MC and BR:

```
R(config-oer-br)# port PORT-NUMBER
```

To enable OER to generate syslog messages, use:

```
R(config-oer-br)# logging
```

To stop the OER process, without deleting the configuration, use:

```
R(config-oer-br)# shutdown
```

### Master Controller

To define a router as the MC, just use:

```
R(config)# oer master
```

Here are a few general commands used on the MC:

```
R(config-oer-mc)# port PORT-NUMBER
! See above. It has to match the same command on the BR
```

The MC and BRs exchange keepalives at a rate that is configurable (5 sec by default). If 3 keepalives are missed, the MC considers the BR down

```
R(config-oer-mc)#keepalive SEC
! Defualt: 5
```

To enable OER to generate syslog messages, use:

```
R(config-oer-mc)# logging
```

To stop the OER process, without deleting the configuration, use:

```
R(config-oer-mc)# shutdown
```

The total number of prefixes that OER can monitor can be customized with the command:

```
R(config-oer-mc)#max prefix total TOTAL learn LEARNED
! Default: TOTAL=5000, LEARNED=2500
```

To verify the status of the MC, use:

```
R# show oer master
! Major Version must match between MC and BR
! Master controller must have a higher or equal minor version
```

#### **Defining BRs**

In the **oer-mc** config mode you can then define the BRs that the MC will control:

```
R(config-oer-mc)# border BORDER-IP key-chain PFR-KEY
! The session is always initiated by the Border Router, so it doesn't need a local interface to be defined
```

Once you define the BR on the MC, you can then select which interfaces on the BR will be part of the PfR setup. Over all BRs you will need to have at least 1 internal and 2 external interfaces.

#### **Internal Interfaces**

An internal interface is used for passive monitoring (using netflow) and normally should be placed on the internal network between the MC and the BRs. To define an internal interface, use:

```
R(config-oer-mc-br)# interface IN-INTERFACE internal
! Config done on the MC
```

#### **External Interfaces**

An external interface is used to forward the traffic and the performance is actively monitored (using IP SLA) by the MC. To define an external interface, use:

```
R(config-oer-mc-br)# interface EX-INTERFACE external
```

For each external interface you can add additional configuration:

```
R(config-oer-mc-br-if)# cost-minimzation ...
! Used to configure the actual cost of routing on one exit interface (through one ISP).
R(config-oer-mc-br-if)# link-group GROUP1 [GROUP2 [GROUP3]]
! Used to add the interface to a link group, which can later be used as a preferred exit point
! An interface can be part of up to 3 Link Groups.
R(config-oer-mc-br-if)# max-xmit-utilization {absolut KBPS|percentage PERCENT}
! The MC will forward traffic on this link as long as the utilization is below the defined value.
! Default: 75%
R(config-oer-mc-br-if)# maximum utilization receive {absolut KBPS|percentage PERCENT}
! Defines the maximum allowed received traffic. Default 75%
R(config-oer-mc-br-if)# downgrade bgp community COMMUNITY
! Adds a BGP community to updates advertising this link in order to downgrade it's usage.
```

Under the global MC configuration mode, you can also define a parameter that is used for proper load-balancing.

```
R(config-oer-mc)# max-range-utilization percent PERCENT
! Default: 20
```

The MC will look at the utilization value for all exit interfaces and if the difference between highest and lowest is more than the above configured parameter, than it considers that the links are not properly balanced so it will try to move some of the traffic from one exit to another.\
There is a similar parmeter for incomint utilization:

```
R(config-oer-mc)# max range receive percent PERCENT
! Default: 20
```

## PfR Policies

PfR uses for its operation a set of rules, called a policy. There is a default policy that is configurable from withing the oer master configuration mode, but you can create your own policies using oer-maps, which are somewhat similar to route-maps or policy-maps.

```
R(config)# oer-map PFR-POLICY [seq SEQ]
```

Inside each entry (each SEQ) you can have **only one match command** and several set commands that override the defaults. The match commands are used to filter the type of traffic that this policy applies to. The filtering can be statically configured or learned dynamically. See [Learning](broken://pages/XGueZUEeHZgBW62HQE3j).

```
R(config-oer-map)# match  ?
  ip             IP specific information
  oer learn      Match dynamically learned OER prefixes
  traffic-class  Specify Traffic class
```

After matching the traffic, you define the actual policy in a series of set commands that override the values of the default policy:

```
R(config-oer-map)# set ...
```

In the end, apply the policy on the MC:

```
R(config-oer-mc)#policy-rules PFR-POLICY
```

You can only apply a single policy-map. If you run the command again with a new policy, it will use the new policy instead of the old policy.\
You can see the configured OER polices using:

```
R# show oer master policy [default|PFR-POLICY]
```

## PfR Stages

PfR uses 5 stages:

1. **Learning (aka Profiling)**: The MC tells the BR to learn the Traffic Classes. These classes can be dynamically learned (via NetFlow) or statically learned (manual configuration).
2. **Measuring Performance**: The BRs collect Traffic Class statistics
3. **Apply Policies**: Use measurements to determine whether a Traffic Class is out of policy (OOP) and if an alternate path can be found.
4. **Enforcing (traffic re-routing)**: Enforcing is done by injecting BGP or static routes, or adding Policy Based Routing (PBR) configuration.
5. **Verification**: Verify that the new route match the policy

### Learning

#### **Static Learning**

For static configuration use one of the following commands:

```
R(config-oer-map)# match ip address {access-list ACL | prefix-list PREFIX}
R(config-oer-map)# match traffic-class application APPLICATION
R(config-oer-map)# match traffic-class {access-list ACL | prefix-list PREFIX}
```

#### **Dynamic Learning**

To configure the policy to use dynamically learned traffic classes, use:

```
R(config-oer-map)# match oer learn {delay|throughput|inside|list LEARN-LIST}
```

Now, you will have to define how learning actually works inside the oer master configuration mode.\
First, to enable dynamic learning, use the following command on the oer master:

```
R(config-oer-mc)# learn
```

Now you can enable learning based on Top Talkers (throughput) and Top Delay

```
R(config-oer-mc-learn)# delay
! Enables learning by Top Delay
R(config-oer-mc-learn)# throughput
! Enable learning by Top Talkers (throughput)
R(config-oer-mc-learn)# inside bgp
! Learns inside prefixes advertised through BGP 
```

Additional learning settings can be configured:

```
R(config-oer-mc-learn)# ?
  aggregation-type   Type of prefix to aggregate. Default: prefix-len 24
  expire             Set expiry criteria for learned prefixes. Default: 720 min
  monitor-period     Period to monitor prefix for learning. Default: 5 min
  periodic-interval  Interval before learning restarts. Default: 120 min
  prefixes           Number of prefixes to learn. Default: 100
```

By default, the MC monitors for 5 minutes and then pauses for 120 min, meaning no routes change for another 120 min. To see results quicker, lower (drastically) these values.\
Traffic lists can be further defined to only learn and categorize specific types of traffic:

```
R(config-oer-mc-learn)# list seq SEQ refname LEARN-LIST
Next, define the traffic classes (more exactly what exactly makes this Traffic Class):
R(config-oer-mc-learn-list)#traffic-class {access-list ACL ...|application ...|prefix-list PREFIX-LIST}
For each least you can specify the method of learning:
R(config-oer-mc-learn-list)#{delay|throughput|rsvp}
You can see the values currently used for Learning with:
R# show oer master | se Learn
				
```

### Measuring

Monitoring the policy can be done passively, actively or using both passive and active monitoring. Passive mode uses Netflow informatation from the BRs (No actual netflow config is needed). Active monitoring uses IP SLA probes that are configured on the MC, but are sourced from the BRs. The other options (fast, both, active throughput) use both active and passive monitoring. To configure the monitoring mode, use:

```
! In the Default Policy:
R(config-oer-mc)# mode monitor {active [throughput] | passive | fast | both}
! In a custome OER-MAP:
R(config-oer-map)# set mode monitor {active [throughput] | passive | fast | both}
To define SLA probes:
! In the Default Policy:
R(config-oer-mc)# active-probe {echo|jitter|tcp-conn|udp-echp} OPTIONS
R(config-oer-mc)# probe frequency SEC
! Default: 60
! In a custom OER-MAP:
R(config-oer-map)# set active-probe {echo|jitter|tcp-conn|udp-echp} OPTIONS
R(config-oer-map)# set probe frequency SEC
! Default: 60
To define the frequency of traceroute probes, use:
! In the Default Policy:
R(config-oer-mc)# traceroute probe delay MSEC
! Default: 1000 msec = 1 sec
! In a custom OER-MAP:
R(config-oer-map)# set traceroute probe delay MSEC
! Default: 1000 msec = 1 sec
```

### Applying Policy

In this stage, OER compares the measurements with thresholds defined in the policy in use. If the perfromance doesn't fit the requirements, OER considers the traffic to be OOPOLICY (Out of Policy) and makes a decision to move it on another exit. To configure the thresholds, use:

```
! In the Default Policy:
R(config-oer-mc)# {delay|loss|jitter|mos|unreachable} [threshold MIN| relative PERCENTAGE] ...
! In a custom OER-MAP:
R(config-oer-map)# set {delay|loss|jitter|mos|unreachable} [threshold MIN| relative PERCENTAGE] ...
The router will try to find a new exit interface for a Traffic Class if it goes OOPOLICY, but also after a fixed period, regardless of the status. This timer is disabled by default but it can be set with:
! In the Default Policy:
R(config-oer-mc)# period SEC
! In a custom OER-MAP:
R(config-oer-map)# set period SEC
When a Traffic Class is moved to a new exit, it will not be changed again until a holddown timer expires. This timer is configured with:
! In the Default Policy:
R(config-oer-mc)# holddown SEC
! Default: 300
! In a custom OER-MAP:
R(config-oer-map)# set holddown SEC
! Default: 300 
Once a Traffic Class becomes OOPOLICY it has to wait a backoff time before beeing moved to an INPOLICY state. This backoff time starts at the MIN value and increases with STEP value for each time the MC tries to move it INPOLICY, but fails. It can't grow more than MAX.
! In the Default Policy:
R(config-oer-mc)# backoff MIN MAX STEP
! Default: MIN=300, MAX=3000, STEP=300
! In a custom OER-MAP:
R(config-oer-map)# backoff MIN MAX STEP
! Default: MIN=300, MAX=3000, STEP=300
When chosing an INPOLICY exit there may be multiple available choices. The MC will chose it based on the order of the resolvers (based on priority. Lowest goes first). Priority 0 is always reachability, so traffic will not be blackholed. To define the order of the resolvers, use:
! In the Default Policy:
R(config-oer-mc)# resolve RESOLVER priority PRI [variance VAR]
! Available Resolvers:
!  cost                         Specify PfR cost policy resolver settings
!  delay                        Specify PfR delay policy resolver setting
!  jitter                       Specify PfR jitter policy resolver settings
!  loss                         Specify PfR loss policy resolver settings
!  mos                          Specify PfR MOS policy resolver settings
!  range                        Specify PfR range policy resolver settings
!  utilization                  Specify PfR utilization policy resolver settings
R(config-oer-map)# set resolve RESOLVER priority PRI [variance VAR]
By default, after Reachability (priority 0), the MC uses Delay (priority 11) and utilization (priority 12) to solve the problem.  The Variance value is used to consider equal multiple values within the same variance. (E.g. Delay values of 300 and 320 are considered equal if the variance is at least 20).
Then, if there are multiple exits that would move the Traffic Class INPOLICY, the router can select the best exit, or any of the good enough exits, based on this command:
! In the Default Policy:
R(config-oer-mc)# mode select-exit {good|best}
! In a custom OER-MAP:
R(config-oer-map)# select-exit {good|best}
! Default: good
When selecting the exit, routers can use by default any external interface. But for some types of traffic, you can select the exit interface on a subgroup of the external interface. Remember that each external interface can be part of up to 3 LINK-GROUPS. If GROUP1 is not available, then GROUP2 can be used.
R(config-oer-map)# set link-group GROUP1 [fallback GROUP2]
For some types of traffic you can select a more static exit strategy:
R(config-oer-map)# set next-hop NEXT-HOP-IP
To drop traffic, just use:
R(config-oer-map)# set interface Null0
```

### Enforcing

The default mode of operation of OER/PfR is "observe mode". In this mode, the MC does what it would normally do, except sending commands to the BRs. This mode is ususually used to see how OER would perform without doing any actual modifications to the traffic paths. To change the mode of operation, use:

```
! In the Default Policy:
R(config-oer-mc)# mode route [control|observe]
! In a custom OER-MAP:
R(config-oer-map)# set mode route [control|observe]
! Default: Observe
```

For traffic classes that are defined using a prefix only, the prefix reachability information used in traditional routing can be manipulated. Protocols such as Border Gateway Protocol (BGP) or RIP are used to announce or remove the prefix reachability information by introducing or deleting a route and its appropriate cost metrics. For traffic classes that are defined by an application in which a prefix and additional packets matching criteria are specified, OER cannot employ traditional routing protocols because routing protocols communicate the reachability of the prefix only and the control becomes device specific and not network specific. This device specific control is implemented by OER using policy-based routing (PBR) functionality. If the traffic in this scenario has to be routed out to a different device, the remote border router should be a single hop away or a tunnel interface that makes the remote border router look like a single hop. PfR uses the concept of parent Route. It won't route traffic on a link unless you have a route for it using that link. In older version of IOS (before 12.4(24)), PFR supported as parent routes only BGP and static routes (even those static floating routes that don't make it into the routing table). With newer versions that support PIRO (Performance Routing Protocol Independent Route Optimization), it should be able to use any type of route

### Verification

PfR uses NetFlow to automatically verify the route control. The MC expects a Netflow update for the traffic class from the new link interface and ignores Netflow updates from the previous path. If a Netflow update does not appear after two minutes, the master controller moves the traffic class into the default state. A traffic class is placed in the default state when it is not under PfR control.


# ODR

## What is ODR?

**ODR**(On Demand Routing) is a routing protocol that runs on top of CDP and that is best suited for a hub and spoke topology. If the CDP traffic is already active between the routers on a **Hub and Spoke** network, the use of ODR makes sense because the connections will not have to transport routing updates for an IGP routing protocol, relying instead on the CDP packets to carry routing information.

Of course, in terms of features, ODR can’t be compared to other more known IGPs like OSPF, EIGRP or even RIP, but in a simple hub and spoke topology, it can make wonders. The key is to be able to transport CDP frames from the hub to every spoke. This is normally the case with Frame Relay or MPLS, but not whithin a Switched Ethernet environment, where CDP would only work on a hop-by-hop basis, so the router’s will see the switches as neighbors, not each other. But there are situations where ODR can work even over ethernet switches:

1. When using dumb switches – They will not intercept CDP packets, but will flood them out all ports.
2. With Cisco swithces (and not just Cisco), using a feature that supports CDP tunneling, like 802.1Q (Q-in-Q) tunneling

Configuring ODR is easy beacuase ODR is enabled ONLY on the hub router with the command:

```
R1(config)# router odr
```

The hub router will then add routing information to the CDP packets it sends to the other spokes, informing them of a default route pointing towards itself. The nice thing is that on the spokes you don’t have to do anything and the routes will be auto-magically set. What can you do to stop the ODR default route from being installed in the spoke’s routing table?

1. You can disable CDP

   ```
   !On an interface:
   R(config-if)# no cdp enable
   ! or globally:
   R(config)# no cdp run
   ```
2. You can start another routing protocol. When this happens, the router will stop processing and sending ODR updates!
3. You can even start ODR on the spoke.\
   When this happens, the spoke will act just like the hub in terms of processing routing information. It will send default routes toward other spokes with himself as next hop but will not exchange routes with other routers configured for ODR. Of course, this will only work if your topology is more of a “partial mesh” because there are practically 2 hubs and the spokes should connect to both of them directly so they can talk over CDP. The spokes will install default routes from each router and this way they will be able to load-balance traffic between the 2 hubs.

ODR has an adminitrative distance of 160 making it less preferred compared to most other protocols. In fact, only External EIGRP routes (AD=170) and Internal BGP(AD=200) are less preferred than ODR.\
ODR supports VLSM and will advertise to the hub all directly connected networks, but only based on the primary IP addresses, not the secondary addresses. ODR will be able to redistribute routes to and from other routing protocols on the hub and it can even do route filtering with the classic command:

```
R(config-rotuer)# distribute-list ACL {in|out} INTERFACE
```

## Just an example

Let’s look now at an example of ODR used in a Frame Relay Topology

ODR Example topology

In this example we will use R1 as the Hub with R2 and R3 as the spokes. R1 and R2 will use the physical interfaces while R3 will use a point-to-point subinterface to the Frame Relay cloud. Each router will have a loopback interface and we will configure ODR on R1 to enable connectivity between each router.

#### 2.1 Basic Frame Relay configuration

```
! Basic config on R1 (HUB):
R1(config)# interface Serial1/1
R1(config-if)# ip address 123.0.0.1 255.255.255.0
R1(config-if)# encapsulation frame-relay
R1(config-if)# no shut
R1(config-if)# exit
R1(config)# interface Lo0
R1(config-if)# ip address 1.1.1.1 255.255.255.255
! Basic config on R2 (SPOKE)
R2(config)# interface Serial1/2
R2(config-if)# ip address 123.0.0.2 255.255.255.0
R2(config-if)# encapsulation frame-relay
R2(config-if)# no shut
R2(config-if)# exit
R2(config)# interface Lo0
R2(config-if)# ip address 2.2.2.2 255.255.255.255
! Basic config on R3 (SPOKE)
R3(config)# interface Serial1/3
R3(config-if)# encapsulation frame-relay
R3(config-if)# no shut
R3(config-if)# interface serial1/3.301
R3(config-subif)# ip address 123.0.0.1 255.255.255.0
R3(config-subif)# frame-relay interface-dlci 301
R3(config-if)# exit
R3(config)# interface Lo0
R3(config-if)# ip address 3.3.3.3 255.255.255.255
```

After LMI and Inverse ARP information is exchanged we end up with:

```
! On R1:
R1#sh frame-relay map
Serial1/1 (up): ip 123.0.0.2 dlci 102(0x66,0x1860), dynamic,
              broadcast,, status defined, active
Serial1/1 (up): ip 123.0.0.3 dlci 103(0x67,0x1870), dynamic,
              broadcast,, status defined, active
! On R2:
R2#sh frame-relay map
Serial1/2 (up): ip 123.0.0.1 dlci 201(0xC9,0x3090), dynamic,
              broadcast,, status defined, active
! On R3:
R3#sh frame-relay map
Serial1/3.301 (up): point-to-point dlci, dlci 301(0x12D,0x48D0), broadcast
          status defined, active
```

#### 2.2 Is CDP working?

In order to make ODR run, we should check if CDP is running:

```
! On R1
R1#sh cdp interface | i Serial
Serial1/0 is administratively down, line protocol is down
Serial1/2 is administratively down, line protocol is down
Serial1/3 is administratively down, line protocol is down
! On R2
R2#sh cdp interface | i Serial
Serial1/0 is administratively down, line protocol is down
Serial1/1 is administratively down, line protocol is down
Serial1/3 is administratively down, line protocol is down
! On R3
R3#sh cdp interface | i Serial
Serial1/0 is administratively down, line protocol is down
Serial1/1 is administratively down, line protocol is down
Serial1/2 is administratively down, line protocol is down
Serial1/3.301 is up, line protocol is up
```

Notice that interfaces Serial1/1 on R1 and Serial1/2 on R2 do not run CDP, while the subinteface Serial1/3.301 on R3 runs CDP! This is because, when using Frame Relay, CDP is enabled by default **only** on point-to-point subinterfaces. We will need to activate CDP now on both R1 and R2:

```
! On R1:
R1(config)# interface s1/1
R1(config-if)# cdp enable
! On R2
R2(config)# interface s1/2
R2(config-if)# cdp enable
```

Let’s verify CDP neighbors now:

```
! On R1:
R1#sh cdp neighbors
Capability Codes: R - Router, T - Trans Bridge, B - Source Route Bridge
                  S - Switch, H - Host, I - IGMP, r - Repeater

Device ID        Local Intrfce     Holdtme    Capability  Platform  Port ID
R2               Ser 1/1            128        R S I      3725      Ser 1/2
R3               Ser 1/1            169        R S I      3725      Ser 1/3.301
! On R2:
R2#sh cdp nei
Capability Codes: R - Router, T - Trans Bridge, B - Source Route Bridge
                  S - Switch, H - Host, I - IGMP, r - Repeater

Device ID        Local Intrfce     Holdtme    Capability  Platform  Port ID
R1               Ser 1/2            124        R S I      3725      Ser 1/1
! On R3:
R3#sh cdp nei
Capability Codes: R - Router, T - Trans Bridge, B - Source Route Bridge
                  S - Switch, H - Host, I - IGMP, r - Repeater

Device ID        Local Intrfce     Holdtme    Capability  Platform  Port ID
R1               Ser 1/3.301        166        R S I      3725      Ser 1/1
```

#### 2.3 Let’s fire up ODR

Now that we have CDP working, we should enable ODR on the hub:

```
R1(config)# router odr
```

That’s it! Let’s check the routing table now:

\[cisco highlite=”7,9,20,29″]! On R1:\
Gateway of last resort is not set

1.0.0.0/32 is subnetted, 1 subnets\
C 1.1.1.1 is directly connected, Loopback0\
2.0.0.0/32 is subnetted, 1 subnets\
o 2.2.2.2 \[160/1] via 123.0.0.2, 00:00:57, Serial1/1\
3.0.0.0/32 is subnetted, 1 subnets\
o 3.3.3.3 \[160/1] via 123.0.0.3, 00:00:16, Serial1/1\
123.0.0.0/24 is subnetted, 1 subnets\
C 123.0.0.0 is directly connected, Serial1/1\
! On R2:\
R2#sh ip route\
Gateway of last resort is 123.0.0.1 to network 0.0.0.0

2.0.0.0/32 is subnetted, 1 subnets\
C 2.2.2.2 is directly connected, Loopback0\
123.0.0.0/24 is subnetted, 1 subnets\
C 123.0.0.0 is directly connected, Serial1/2\
o\* 0.0.0.0/0 \[160/1] via 123.0.0.1, 00:00:37, Serial1/2\
! On R3:\
R3#sh ip route\
Gateway of last resort is 123.0.0.1 to network 0.0.0.0

3.0.0.0/32 is subnetted, 1 subnets\
C 3.3.3.3 is directly connected, Loopback0\
123.0.0.0/24 is subnetted, 1 subnets\
C 123.0.0.0 is directly connected, Serial1/3.301\
o\* 0.0.0.0/0 \[160/1] via 123.0.0.1, 00:00:23, Serial1/3.301

Well, it looks pretty good. We have default routes on the spokes and the hub knows how to get to all networks. Now let’s test this, just to make sure:

```
!On R1:
R1#ping 2.2.2.2

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 2.2.2.2, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 28/36/40 ms
R1#ping 3.3.3.3

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 20/40/68 ms
!On R2:
R2#ping 1.1.1.1

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 1.1.1.1, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 24/33/40 ms
R2#ping 3.3.3.3

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 3.3.3.3, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 56/77/88 ms
!On R3:
R3#ping 1.1.1.1

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 1.1.1.1, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 24/36/48 ms
R3#ping 2.2.2.2

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 2.2.2.2, timeout is 2 seconds:
.....
Success rate is 0 percent (0/5)
```

Wow! It all seemed to go well until the last ping. What happend? How come we can ping R3 from R2 but not R2 from R3?

#### 2.4 What went wrong?

Well, this doesn’t really have to do with ODR anymore. ODR did its job – we have correct routing information on all routers.\
What happens is that when we ping R3 from R2, we actually send packets with source 123.0.0.2 and destination 3.3.3.3. R2 will look in its routing table and will find that is should send packets for 3.3.3.3 using the default route. The default route points to 123.0.0.1, and we have an Inverse ARP map for that – 102. The packet is sent out to R1. On R1, it will be redirected to R3. When R3 receives the ICMP packet, in order to reply, it will create a packet with a source of 3.3.3.3 and destination of 123.0.0.2.

**Now here’s why R2 can ping R3’s loopback**: To send to 123.0.0.2 the routing table points to the directly connected interface Serial1/3.301. This is a point-to-point interface and R3 doesn’t care about Inverse ARP or static maps on this interface. It will just send the packet using the DLCI assigned to it. THe packet arrives at R3 who redirects it to R2 and we have connectivity!

In the same way, when we ping R2 from R3, we actually send packets with source 123.0.0.3 and destination 2.2.2.2. R3 looks up a route to 2.2.2.2 It doesn’t find one so it uses the default route that points to 123.0.0.1 on the point-to-point Serial1/3.301. R3 just encapsulates the packet using DLCI 301 and sends it out. R1 receives the packet and redirects it over to R2.

**Now here’s why R3 can’t ping R2’s loopback**: R2 receive the packet and must send a reply from source ip 2.2.2.2 to destination IP 123.0.0.3. It looks up the destination in the routing table, finds out it’s Serial1/2 (directly connected) but we don’t have a mapping for 123.0.0.3 and we don’t know what DLCI to use to reach this IP. Encapsulation fails, the ping fails.

#### 2.5 Solutions?

The solution? We can just ping R2’s loopback using R3’s loopback. There’s no problem with this

```
R3#ping 2.2.2.2 source 3.3.3.3

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 2.2.2.2, timeout is 2 seconds:
Packet sent with a source address of 3.3.3.3
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 64/76/92 ms
```

Or, we can add a static map on R2 that tells the router to use the same DLCI 201 to reach 123.0.0.3, also.

```
! On R2:
R2(config)#int s1/2
R2(config-if)#frame-relay map ip 123.0.0.3 201
! On R3:
R3#ping 2.2.2.2

Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 2.2.2.2, timeout is 2 seconds:
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 64/81/92 ms

3. Conclusions
ODR works great, but there's a difference between having a route and having connectivity. Especially on Frame Relay.
							
					
```


# IPv6-101

## IPv6 Packet Format

The IPv6 header has a fixed format (as opposed to the variable format of IPv4) of 40 Bytes

![IPv6 Header](/files/acbHXZiTRy4vQcz4Cc5m)

**Version** – 4 bits – always set to the value 6

* **Traffic Class** – 8 bits – 6 most significat bits are used for [DSCP](https://nyquist.eu/classification-and-marking/#222_DSCP), 2 least significant bits are used as [ECN](https://nyquist.eu/congestion-avoidance-wred/#221_WRED_Explicit_Congestion_Notification_ECN)
* **Flow Label** – 20 bits – Usage not completly standardized, but usually used to mark packets that should follow the same path in a multi-path environment.
* **Payload length** – 16 bits – size of the payload in Bytes, including any extension headers
* **Next Header** – 8 bits – Specifies the next header – either an extension header or the next layer (usually TCP) header
* **Hop Limit** – 8 bits – Similar to TTL
* **Source Address** – 128 bits
* **Destination Address** – 128 bits

## IPv6 Address Format

One of the features of IPv6 is the big address space. This is achieved by using addresses 128 bits long. IPv4 addresses were only 32 bits long.\
Such a big address number becomes difficult to represent in a human readable format, and the old convention used for IPv4 (dotted decimal: A.B.C.D) cannot be used anymore. The format that is used for IPv6 is of 8 groups of 16 bits, written in hex. An example could be:

```
2001:0DB8:0000:0000:0008:0000:0000:417A
```

To use an even shorter notations, 2 new rules are used:

1. In any of the 8 groups, leading zeros can be omitted, but if they are all zeros, one stil has to show up. Our example becomes:

   ```
   2001:0DB8:0:0:8:0:0:417A
   ```
2. Once, replace one or more groups of zeros with “::”. Our example becomes:

   ```
   2001:0DB8::8:0:0:417A
   !or 
   2001:0DB8:0:0:8::417A
   ```

### EUI64

EUI64 is a method of generating a unique Interface ID. [RFC 2464](https://tools.ietf.org/html/rfc2464) shows how to generate a unique Interface ID from the MAC Address of the interface.\
Split the MAC address in 2 equal parts, insert FF:FE in the middle to reach the required 64 bits and flip the U bit (Universal/Local bit, the 7th bit in the first Byte). It ends up subtracting (or adding) 2 from the second HEX digit of the MAC Address. If the interface doesn’t have a MAC address, a router may use a MAC addresses assigned to the router, the Serial Number, an md5 hash of the hostname or a random number to generate the EUI-64 address.\
Example

![EUI64 method](/files/XYmXbclc6uo68WdQiP77)

## IPv6 Address Types

IPv6 address usage has changed over the years and some types are now deprecated. A current address space and the address types can be seen at [IANA’s site](https://www.iana.org/assignments/ipv6-address-space/ipv6-address-space.xml). The current IETF RFC that deals with the Addressing Architecture is [RFC 4291](https://tools.ietf.org/html/rfc4291).

### Unicast

On an interface you can have multiple IPv6 addresses. There is no “secondary” address, all of them are “primary”. [RFC 3484](https://www.ietf.org/rfc/rfc3484.txt) describes what address should be used to source traffic.

You can set unicast address manually, using:

```
R(config-if)# ipv6 address IPV6-ADDRESS/PREFIX-LEN
! Using EUI-64:
R(config-if)# ipv6 address IPV6-PREFIX/PREFIX-LEN eui-64
! Using another interface IPv6 address
R(config-if)# ipv6 address unnumbered INTERFACE
```

Unicast addresses can also be dynamically assigned using DHCPv6 or autoconfig. Autoconfig will use eui-64 for the Interface ID and the prefix received in RA from the router:

```
R(config-if)# ipv6 address autoconfig [default]
! default - will also insert a default route in the routing table
```

#### **Global Unicast 2000::/3**

The format is

```
| 3 |        45 bits        |  16 bits  |    64 bits   |
+---+-----------------------+-----------+--------------+
|001| global routing prefix | subnet ID | Interface ID |
+---+-----------------------+-----------+--------------+
```

Address space: 2000:: –> 3FFF:…:FFFF.\
The interface ID can be represented in EUI-64 format based on the MAC address or it can be manually assigned.

#### **Link Local Unicast FE80::/10**

Link Local addresses are used on a single link (point-to-point or multi-access) and are used for autoconfiguration, neighbor discovery, and so on. They are not forwarded out of their scope. The format used for Link Local addresses is:

```
|  10 bits |   54 bits |       64 bits      |
+----------+-----------+--------------------+
|1111111010|   00..0   |    Interface ID    |
+----------+-----------+--------------------+
```

Address space: FE80:: –> FEBF:…:FFFF.\
By default, when an interface comes up, it automatically generates a link-local address using the FE80::/10 prefix and the EUI64 Interface ID. To override the automaticly generated link-local address, use:

```
R(config-if)# ipv6 address IPV6-ADDRESS link-local
```

Pinging a Link Local Unicast Address requires declaring what interface to use.

#### **IPv4 compatibile**

An IPv4 compatibile IPv6 address contains 96 bits of zero followed by 32 bits of the IPv4 address:

```
|  96 bits |      32 bits    |
+----------+-----------------+
|  00..0   |  IPv4 address   |
+----------+-----------------+
```

#### **Unique Local Address (ULA) FC00::/7**

Unique Local addresses are standardized by [RFC 4193](https://tools.ietf.org/html/rfc4193) and are intended for local use, not to be routed in the Internet. They are similar to IPv4 local addresses (10.0.0.0/8, 172.16.0.0/12 and 192.168.0.0/16). The format is:

```
| 7 bits|L|   40 bits   |  16 bits  |    64 bits   |
+-------+-+-------------+-----------+--------------+
|1111110|1|  Global ID  | subnet ID | Interface ID |
+-------+-+-------------+-----------+--------------+
```

Address space: FC00:: –> FD00:…FFFF. Since only addresses with the 8th bit set to 1 are permitted, this actually means the usable space is FD00::/8.

### Anycast

An anycast address is an address that is assigned to multiple interfaces. The difference between anycast and multicast is that while a packet sent to a multicast address will reach all interfaces in the multicast group, a packet sent to an anycast address will reach only the interface that is “closest” in terms of routing. This is useful in representing geographically different hosts with the same IP address.\
On a subnet, each router must be able to respond to the Anycast address with all-zeros in the host field. This is the Subnet-Router Anycast Address:

```
|      n bits        |      128-n bits      |
+--------------------+----------------------+
|    subnet prefix   |         00..0        |
+--------------------+----------------------+
```

Anycast addresses are set similar to a global unicast address, but the **anycast** keyword must be used:

```
R(config-if)# ipv6 address IPV6-ADDRESS/IPV6-PREFIX-LEN anycast
```

### Mutlicast FF00::/8

See [IPv6 Multicast](https://nyquist.eu/multicast-ipv6)

### Unspecified and Loopback Addresses

The unspecified address is an all-zeros address, also written as ::/128 and indicates no IPv6 address assigned on a specific interface.\
The loopback address is used by a host to send packets to itself. It has 127 zeros and one last bit of 1. It is written as ::1/128

## Neighbor Discovery

ND, defined in [RF2461](https://tools.ietf.org/html/rfc2461), is a process that uses ICMPv6 messages to replicate and ehance IPv4 ARP features. ND defines 5 ICMP packet types:

* **Router Solicitation – RS**: When an interface comes up, a RS is sent to request an RA from the router
* **Router Advertisement – RA**: Packets used by routers to advertise their presence. They are sent periodically or as a response to a Router Solicitation packet
* **Neighbor Solicitation – NS**: Sent by a node to determine the Link Layer address (MAC on Ethernet) of a neighbor. Also used for Duplicate Address Detection
* **Neighbor Advertisement – NA**: Sent in response to a NS. A node can also send unsolicited NA when its link-layer addres (MAC on Ethernet) changes
* **Redirect**: Used by routers to inform hosts of a better next-hop for a destination

These messages are used to offer the following features:

* **Router Discovery**: Hosts locate the routers on their link
* **Prefix Discovery**: Hosts discover their prefix on the link
* **Parameter Discovery**: Hosts discover other parameters, like MTU or TTL of outgoing packets
* **Address Auto configuration**: Hosts will autoconfigure an address based on the link-local prefix or the prefix advertised by the router and the EUI-64 Interface ID.
* **Address Resolution**: Finding a neighbor’s Layer 2 address (similar to IPv4 ARP)
* **Next-hop determination**
* **Neighbor Unreachability Detection**: NS and ND messages are sent in order to verify that a neighbor is reachable or not.
* **Duplicate Address Detection**
* **Redirect**

On a router interface you can tune ND parameters, using:

```
R(config-if)# ipv6 nd OPTIONS
```

To see IPv6 neighbors (from NA messages, similar to IPv5 ARP cache), use:

```
R# show ipv6 neighbors
```

To see IPv6 routers from (from RA messages), use:

```
R# show ipv6 routers
```

To see how Neighbor Discovery works you can enable debugging with:

```
R# debug ipv6 nd
```


# IPv6 Routing

## Enable IPv6 Routing

To enable IPv6 routing, use:

```
R(config)# ipv6 unicast-routing
```

This will enable the router to send RAs unsolicited or in response to RS messages.

## Static Routing

```
! Directly attached static routes:
R(config)# ipv6 route IPV6-PREFIX/IPV6-PREFIX-LEN OUT-INTERFACE [AD]
! Recursive routes:
R(config)# ipv6 route IPV6-PREFIX/IPV6-PREFIX-LEN IPV6-NEXT-HOP [AD]
! Fully Specified Static routes:
R(config)# ipv6 route IPV6-PREFIX/IPV6-PREFIX-LEN INTERFACE IPV6-NEXT-HOP [AD]
```

To see the routing table, use:

```
R# show ipv6 route
```

## RIP for IPv6

RIP for IPv6, aka RIPng, works just like RIPv2 for IPv4. It sends multicast packets to FF02::9 multicast address.

### Start the process

To start the RIP process, enable RIP on each interface where it should run:

```
R(config-if)# ipv6 rip PROC-NAME enable
```

RIP uses the link-local addresses as RIP source address, so in Frame Relay you would probably need static mapping. There can be multiple instances of RIPng running on the same router and you can set each to listen on a different UDP port on the same subnet.

```
R(config)# ipv6 router rip PROC-NAME
R(config-rtr)# port PORT multicast-group MULTICAST-ADDR
```

### Default routes

To generate a default route use the command:

```
R(config-if)# ipv6 rip PROC-NAME default-information {originate|only} [metric METRIC]
! originate - advertises the default and the other RIP routes
! only - advertises only the default (like a summary)
```

## EIGRP for IPv6

### Start the process

Enable EIGRP on each interface:

```
R(config-if)# ipv6 eigrp AS-NUMBER
```

By default, the EIGRPv6 process is shutdown, so we must enable it with:

```
R(config)# ipv6 router eigrp AS-NUMBER
R(config-rtr)# no shut
```

If there are no IPv4 addresses on the router, it cannot create a Router ID and the process won’t start. You will have to statically define one:

```
R(config-rtr)# router-id IPV4-ADDRESS
```

There is no auto-summary in EIGRP for IPv6.

EIGRP uses the link-local addresses to create adjacencies so you might need to set static mappings in Frame Relay.

EIGRP for IPv6 uses FF02::A address for multicast messages.

There are no network statements in EIGRP for IPv6, only interface statements

## OSPFv3 for IPv6

It works similar to OSFPv2.

### Start the OSPFv3 Process

Assign interfaces to OSPFv3:

```
R(config-if)# ipv6 ospf PROC-ID area AREA-ID
```

If there are no IPv4 addresses on the router, it cannot create a Router ID and the process won’t start. You will have to statically define one:

```
R(config)# ipv6 router ospf PROC-ID
R(config-rtr)# router-id IPV4-ADDRESS
```

OSPF uses the link-local addresses to create adjacencies so you might need to set static mappings in Frame Relay.

### Summarization

```
! at the ABR:
R(config-rtr)# area AREA range IPV6-PREFIX
! at the ASBR
R(config-rtr)# summary-prefix IPV6-PREFIX [not-advertise|tag TAG]
! not-advertise will filter both the summary and the children
```

## MP-BGP for IPv6

IPv6 can be advertised via MP-BGP. It is just another address-family available in the normal BGP process.

### Start MP-BGP for IPv6

```
R(config)# router bgp AS-NUMBER
R(config-router)# neighbor IPV6-NEIGH-ADDR remote-as REMOTE-AS
R(config-router)# address-family ipv6
R(config-router-af)# neighbor IPV6-NEIGH-ADDR activate
```

If there are no IPv4 addresses on the router, it cannot create a Router ID and the process won’t start. You will have to statically define one:

```
R(config-router)# bgp router-id IPV4-ADDRESS
```

## Redistribution

By default with IPv4 protocols, when redistributing from one protocol to another, the connected routes, matched by the network command where also redistributed. With IPv6 routing protocols, they are not redistributed by default anymore. Instead you will have to use the **include-connected** keyword:

```
R(config-rtr)# redistribute PROTOCOL [PROCESS|AS-NUMER] [include-connected] OPTIONS
```

## IPv6 Policy Based Routing

It works similar to [IPv4 PBR](https://nyquist.eu/policy-based-routing/).\
To enable the policy, use:

```
R(config-if)# ipv6 policy route-map ROUTE-MAP
! for locally originated traffic
R(config)# ipv6 local policy ROUTE-MAP
```


# Interconnecting IPv6 and IPv4

## Dual Stack

Dual Stack means running both IPv4 and IPv6 on the same interface. This doesn’t mean they are actually interconnected, just that they can both run at the same time on an interface.

## Tunnels

### Manual Tunnels

#### **IPv6 in IP Tunnels**

IPv6 can be encapsulated in IPv4 packets and transported across a virtual tunnel interface. To enable a tunnel that encapsulates IPv6 in IPv4, use:

```
R(config)# interface TUNNEL0
R(config-if)# tunnel mode ipv6ip
```

Also, a tunnel source and a destination must be set. These must be IPv4 addresses. For tunnel destination you can also configure an interface running IPv4.

```
R(config-if)#tunnel source {INTERFACE|SRC-IP-ADDR}
R(config-if)# tunnel destination DEST-IP-ADDR
```

#### **GRE Tunnels**

GRE Tunnels work the same, except the encapsulation protocol is GRE which works both with IPv4 and IPv6. The advantage with GRE is that it can encapsulate both IP and IPv6 and it can be transported with both IP and IPv6. GRE over IPv4 is the default encapsulation mode of Cisco Tunnels. You can set the tunnel mode to gre with:

```
R(config-if)# tunnel mode gre {ip|ipv6}
! gre ip: GRE over IPv4
! gre ipv6: GRE over IPv6
```

Also, a tunnel source and a destination must be set:

```
R(config-if)#tunnel source {INTERFACE|SRC-IP-ADDR}
R(config-if)# tunnel destination DEST-IP-ADDR
```

The transported protocol depends on the type of address that is set on the tunnel interface. It can be either an IPv4 address or an IPv6 address,

```
R(config-if)# ip address ..
R(config-if)# ipv6 address ..
```

### Automatic Tunnels

Automatic tunnels are inherently multipoint because the destination of the tunnel is dynamic, computed out of the IPv6 destination address of each packet. When an IPv6 packet is routed out the auto-tunnel interface, the router looks at the destination IPv6 which has to follow a certain rule based on the type of the auto tunnel. From the destination IPv6 address, it computes the IPv4 tunnel destination address. Now it know wheat is the tunnel destination for this particular packet. The process is performed again for each packet.\
Things happen the same on the return path, as well.\
You also have to make sure the routing table points IPv6 traffic destined for the tunnel on the correct tunnel interface.

The automatic tunnels are all based on ipv6ip encapsulation, so the configuration is similar:

```
R(config-if)# tunnel mode ipv6ip {6-to-4|auto-tunnel|isatap}
```

#### **6to4 Tunnels**

IANA Reserved the 2002::/16 prefix for 6to4 tunnels. The prefix of the address used should be in the format **2002:HAHB:HCHD::** where HA, HB, HC and HD are the hex values of A, B, C and D in the IPv4 address of the tunnel source: A.B.C.D.

To define the 6to4 automatic tunnel mode, use:

```
R(config)# tunnel TUNNEL0
R(config-if)# tunnel mode ipv6ip 6to4
```

Then, set a tunnel source but do not set a tunnel destination.

```
R(config-if)# tunnel source {INTERFACE|SRC-IP-ADDRESS}
! IPv4 source address is: A.B.C.D.
```

Then, set the IPv6 address of the tunnel interface.

```
R(config-if)# ipv6 address 2002:HAHB:HCHD::1/PREFIX-LEN
! Compute HA, HB, HC, HD based on the A.B.C.D address of the tunnel source
```

Next, a route to the peer(X.Y.Z.T) should be set. This must be done for each potential peer

```
R(config)# ipv6 route 2002:HXHY:HZHT::/PREFIX-LEN TUNNEL0
```

Automatic 6-to-4 tunnels can only be used for BGP peering, which uses unicast addresses. Other routing protocols use Link Local addresses and can’t work over an automatic 6-to-4 tunnel.

#### **Auto-Tunnels**

This feature is deprecated in real life scenarios, but it can be configured in IOS. This is also a point-to-multipoint configuration like 6-to-4 tunnels, so you don’t need to set a tunnel destination. You also don’t have to set an IPv6 address on the tunnel, as it was required in 6-to-4 tunnels.\
To configure auto-tunnels, just set the mode and the tunnel source:

```
R(config)# interface TUNNEL0
R(config-if)# tunnel mode ipv6ip auto-tunnel
R(config-if)# tunnel source {INTERFACE|SRC-IP-ADDRESS}
The auto-tunnel mode will automatically set an IPv6 Address on the tunnel interface, in the format ::HAHB:HCHD or ::A.B.C.D (these notations represent the same address).
Next, if you want to reach networks past the next hop router, a static routs or BGP can be used. A static route could look like:
R(config)# ipv6 route IPV6-PREFIX/PREFIX-LEN ::HXHY:HZHT
! or
R(config)# ipv6 route IPV6-PREFIX/PREFIX-LEN ::X.Y.Z.T					
```

#### ISATAP Tunnels

ISATAP(Intra-Site Automatic Tunnel Addressing Protocol) is also a point-to-multipoint configuration, like the other automatic tunnels. You don't need to setup a tunnel destination, but you need to set the IPv6 tunnel address using the ISATAP format. To configure it:

```
R(config)# interface TUNNEL0
R(config-if)# tunnel mode ipv6ip isatap
R(config-if)# tunnel source {INTERFACE|SRC-IP-ADDRESS}
R(config-if)# ipv6 address PREFIX/64 eui-64
! EUI-64 Interface ID will be in the format:0000:5EFE:HAHB:HCHD 
```

The router will automatically create an ISATAP specific eui-64 Interface ID in the following format: **0000:5EFE:HAHB:HCHD**, where A.B.C.D is the source of the tunnel. The IPv6 address of the tunnel, will therefor be **PREFIX::0000:5EFE:HAHB:HCHD/64** Next, if you want to reach networks past the next hop router, static routes or BGP can be used. A static route could look like:

```
R(config)# ipv6 route IPV6-PREFIX/PREFIX-LEN IPV6-TUN-PREFIX::5EFE:HXHY:HZHT
```

The advantage of using ISATAP is that it also has link-local addresses and therefor we should be able to run IGP protocol on such tunnels. Unfortunately the tunnels cannot use the multicast addresses so you will have to manually specify the neighbors. Another advantage is that the prefix is left to the administration's decision and is not fixed as in the other automatic tunnel modes.

## NAT-PT

NAT PT has been deprecated but Cisco IOS still supports it. It is similar to IPv4 NAT, but it can also translate the protocol from IPv4 to IPv6 and the other way around. To better understand the configuration, we will use this example:&#x20;

![NAT-PT](/files/X8fXT2TB9EkOgv8rgYtg)

&#x20;NAT PT requires a few steps when configuring:

### Basic NAT PT config

Enable NAT PT on each interface and define the NAT PT Prefix:

```
R(config)# ipv6 unicast-routing
R(config)# interface IPV6-INTERFACE
R(config-if)# ipv6 nat
R(config-if)# exit
R(config)# interface IPV4-INTERFACE
R(config-if)# ipv6 nat
R(config-if# exit
```

Enabling NAT-PT will enable a virtual interface (NVI) on the router. For this interface we must define a prefix. This prefix must always be /96. (128-32)

```
R(config)# ipv6 nat prefix NAT-PT-PREFIX/96
```

The next to steps envolve defining a v6v4 and a v4v6 translation. Each translation can be configured using Static or Dynamic NAT. Dynamic NAT has the inconvenient that it can only be used on one side, and that side must initiate the Session. Static NAT however will be difficult to implement in many-to-many communications

### Define a v6v4 translation

The translation can be done using Static NAT or Dynamic NAT.

#### Static NAT

Translates the IPV6-SRC to an IPV4-ADDR on the IPv4 network or another subnet used only for NAT:

```
R(config)# ipv6 nat v6v4 source IPV6-SRC IPV4-ADDR
```

#### Dynamic NAT

Translates any IPv6 Address matched by an ACL or route-map into an IPv4-Address specified in a pool or used on an interface:

```
!using an interface:
R(config)# ipv6 nat v6v4 source {list ACL| route-map ROUTE-MAP} interface INT
! using a POOL
R(config)# ipv6 nat v6v4 source {list ACL| route-map ROUTE-MAP} pool V4-POOL
```

The V4-POOL must be defined using:

```
R(config)# ipv6 nat v6v4 pool V4-POOL IPv4-START IPv4-END prefix-length MASKLEN
```

#### Dynamic NAT with overload - PAT

PAT works the same as in IPv4 NAT. Instead of mapping each source to an IP address of the pool, the router will map multiple IPv6 addresses to the same IPv4 address, changing the L4 source Port in the translation.

```
R(config)# ipv6 nat v6v4 source {list ACL| route-map ROUTE-MAP} interface INT overload
! using a POOL
R(config)# ipv6 nat v6v4 source {list ACL| route-map ROUTE-MAP} pool V4-POOL overload
```

The V4-POOL must be defined using:

```
R(config)# ipv6 nat v6v4 pool V4-POOL IPv4-START IPv4-END prefix-length MASKLEN
```

### Define a v4v6 translation

The final step in making the communication work end-to-end is to define the translation for the traffic coming in the opposite direction. There are also a few options:

#### Static NAT

```
R(config)# ipv6 nat v4v6 source IPV4-SRC IPV6-ADDR
```

#### Dynamic NAT

```
R(config)# ipv6 nat v4v6 pool V6-POOL IPV6-START IPV6-END prefix-length MASKLEN
```

#### Automatic Mapping

You can configure the router to make automatic mappings between IPv4 and IPv6 addresses by adding the 32 bits of the IPv4 address to the 96 bits IPv6 NAT PT prefix and the other way around. To enable this feature, use:

```
R(config)# ipv6 nat prefix NAT-PT-PREFIX/96 v4-mapped ACL
! In our example:
R(config)# ipv6 nat prefix 2003::/96 v4-mapped ACL
```

The ACL is used to filter the traffic that will go through the NAT process.


# MPLS 101

## Labels

MPLS (Multi Protocol Label Switching) adds a 4 Byte header between the Layer 2 and Layer 3 header for each label. The header has 4 fields:

* 20 bits – Label – locally significant on the link (similar to Frame Relay DLCI)
* 3 bits – EXP – used for QoS markings
* 1 bit – Bottom of Stack – signals the last label
* 8 bits – TTL

Certain applications of MPLS require the use of multiple labels. This is achieved by adding multiple MPLS headers before the L3 header. This operation is called “label stacking”. The last label has the Bottom of Stack (BoS) bit set to 1. All others have it set to 0.

Routers participating in MPLS operations are called LSR (Labels Switch Routers). Based on their position in the traffic path, they act as Ingress LSR (receives unlabeled packets and sends labeled packets), Intermediate LSR(receives labeled packets ands send labeled packets – but the label was swapped, as it has only local significance) and Egress LSR (receives labeled packets and send unlabeled packets).

MPLS labels are bound to FECs (Forwarding Equivalency Classes). All packets inside a FEC follow the same route to the next hop, so they are marked with the same label. In a network that doesn’t use traffic engineering, each IGP destination (including static and connected) should represent a FEC.\
The operations allowed on labels are:

* Push – replace the top label with another and add a new label on top.
* Swap – replace the top label with another label
* Pop – remove the top label

### Enable MPLS per interface

CEF must be on:

```
R(config)# ip cef
```

Then, start MPLS on each interface:

```
R(config-if)# mpls ip
```

This command will automatically start LDP or TDP based on the following setting:

```
! Global:
R(config)# mpls label protocol {ldp|tdp}
! Per interface:
R(config-if)# mpls label protocol {ldp|tdp|both}
```

To verify, use:

```
R# show mpls interfaces
R# show mpls ldp neighbors
```

### TTL

By default, the IP TTL field is copied into the MPLS TTL and decremented at each hop. This will make the MPLS cloud visible in traceroutes. To make it invisible:

```
R(config)# no mpls ip propagate-ttl
```

When the packet reaches the Egress router, the MPLS TTL value is copied into the IP TTL field (except when the MPLS TTL value is higher than the existing IP TTL value to prevent loops)\
In a label stack, an LSR only modifies the top label and doesn’t modify the other labels below it.\
When the TTL reaches 1, the TTL expires and the packet is dropped. In classic IP networks, an ICMP “Time exceeded” message would be sent back to the original sender. In MPLS, it can happen that the router that dropped the packet doesn’t know how to reach the source of the packet. To address this, the router sends the ICMP message towards the destination, in order to find the Egress router. The Egress router should know how to reach the source of the dropped packet and it will now send the ICMP packet on the path to the source. (it may very well go through the router that dropped the packet initially).

## Label Distribution

Label bindings can be set static or dynamic. For dynamic distribution of labels between LSRs, you can use:

* TDP (Tag Distribution Protocol – Cisco prioprietary) – similar to LDP
* LDP (Label Distribution Protocol)
* RSVP (Resource reSerVation Protocol) – used for MPLS Traffic Engineering
* MP-BPG(Multi Protocol BGP)

### LDP

#### **LDP Neighbor Adjacency**

LDP uses UDP 646 to 224.0.0.2 for neighbor discovery. Once discovered, it will create a Neighbor Adjacency using TCP 646 to remote LDP Router ID using its own Router-ID as the source. If you don’t want to use the Router-ID as the source, use:

```
R(config-if)# mpls ldp discovery transport-address {IP-ADDR | INTERFACE}
```

The Router ID is detetrmined just like in OSPF and it can be overridden using:

```
R(config)# mpls ldp router-id {IP-ADDR|INTERFACE}
```

To verify LDP adjacencies, use:

```
R# show mpls ldp neighbors
```

#### **Label bindings**

Each router that runs MPLS, will assign labels to all routes in its routing table( connected interfaces and IGP learned routes). These are called local bindings. Usually label allocation start with 16 and move forward. You can see the label allocation space with the command:

```
R#show mpls label range
Downstream Generic label region: Min/Max label: 16/100000
```

Via LDP, or another label distribution protocol, it will advertise the local bindings to its neighbors and will also receive bindings advertised by its peers. These are called remote bindings. The bindings make up the LIB (Label Information Base). You can see it using:

```
R# show mpls ldp bindings [DESTINATION]
!E.g:
R# show mpls ldp bindings 10.9.9.0 24
  tib entry: 10.9.9.0/24, rev 30
       local binding:  tag: 25
        remote binding: tsr: 3.3.3.3:0, tag: 47
        remote binding: tsr: 4.4.4.4:0, tag: 32
        remote binding: tsr: 5.5.5.5:0, tag: 64
```

If a LSR receives binding information for the same destination from multiple LSR neighbors, it will select the information coming from the next hop of that specific destination taken out of the routing table (RIB). The selected information will make up the LFIB (Label Forwarding Information Base) and it will include local label, outgoing operation (pop, swap), destination prefix, outgoing interface and next hop. You can check the LFIB with:

```
R# show mpls forwarding-table [DESTINATION]
!E.g:
R# show mpls forwarding-table 10.9.9.0 24
Local  Outgoing    Prefix            Bytes tag  Outgoing   Next Hop
tag    tag or VC   or Tunnel Id      switched   interface
25     32          10.9.9.0/24       0          Fa0/1     3.3.3.3
26     45          10.8.8.0/24       0          Se0/1/0   point2point
27     Pop Tag     10.7.7.0/24       0          Fa0/1     3.3.3.3
28     Untagged    10.6.6.0/24       0          Fa0/2     6.6.6.6
```

Pay attention to entries that only have a local or a remote binding. This might indicate a problem with the IGP. If an entry only has local entry, it means no other neighbor knows about this prefix. Similarly, if only remote entries appear, then this router doesn’t know how to reach the prefix advertised by one of its LDP routers. Routes that are directly connected are advertised with a NULL label, to let the penultimate hop know that the next router will forward the packet based on the IP address. See [PHP](broken://pages/FVtJxR3mZxY3f8u1wya6) for details.

The operation field in the LFIB will show:

* **Pop tag** – if the top label will be popped. Usually when [PHP](broken://pages/FVtJxR3mZxY3f8u1wya6) is in action
* **LABEL** – for the SWAP operation – Usually for all destinations that are more than 1 hop away
* **Untagged** – when packets are forwarded as IP – for all destinations that point to non MPLS neighbors

For ingress routers, forwarding is done based on CEF entries, and you can see there how a label is imposed (PUSH):

```
R# show ip cef [DESTINATION] 
```

#### **PHP – Penultimate Hop Popping**

PHP is an MPLS feature which says that if a packet must be forwarded to the last MPLS enabled router in the path (the PE), the MPLS label should be popped, and the packet should be sent as IP. The reason for this is that if we would send the packet labeled, the egress router should perform both an MPLS lookup based on the label and an IP lookup based on the destination IP. The routers advertise the destinations that do not run MPLS with an “Implicit NULL” label (Label value: 3). This way, the penultimate router knows that it should POP labels when it is sending for those destinations. This is the default behavior of Cisco IOS routers.

With that in mind, remember that there is an EXP field in the MPLS header that is used for QoS markings. If PHP is in place, the MPLS header is removed and QoS information is lost when the packet reaches the Egress router. To avoid these situations, the egress router advertise the destinations with an “Explicit NULL” label (Label value: 0 for IPv4, 2 for IPv6). This will make the penultimate router to send the packets for that prefix with an MPLS header which includes the QoS markings in the EXP field and NULL in the Label field. To configure an egress router to advertise Explicit NULL labels, use:

```
R(config)# mpls ldp explicit-null [for PREFIX-ACL] [to PEER-ACL]
```

#### **Limit LDP Advertisements**

By default, LDP will advertise labels for all entries in the routing table. You can disable this, by using:

```
R(config)#no mpls ldp advertise-labels
```

and then use a standard ACL to advertise only specific routes:

```
R(config)# mpls ldp advertise-labels for ACL
```


# MPLS L3 VPN

This article assumes the “provider” network already has an IGP in place and that the LDP is configured to advertise label bindings between LSRs. Check [MPLS 101](/mpls/mpls-101) on how to do that.

## Verify LDP is working within provider network

One common mistake when configuring L3 MPLS VPN appears when using OSPF as the IGP. L3 MPLS VPN uses 2 stacked labels when sending customer traffic from one PE to another. The outer most label identifies traffic for the PE and should be popped before reaching the PE, through PHP. When using OSPF, loopback addresses are advertised by default as /32 prefixes, regardless of the actual subnet. However, in MPLS they will be advertised with their subnet. If the loopback address is different then /32 the PE router will advertise via LDP a network with, let’s say, /24 (because the Connected route will be in it’s routing table), while it’s LDP neighbor will advertise a /32 route (known via OSPF). They will not agree upon a label, and the penultimate router will consider that it should route the packets as IP, not as MPLS. You can spot such mistakes when looking in the MPLS forwarding table with:

```
R# show mpls forwarding-table
```

Prefixes that point to the PE’s loopback should have the operation POP. If the operation is UNTAGGED, then there is a mismatch between the advertisement of the PE and the penultimate router.\
Usually, in a working MPLS network, you should only see UNTAGGED on egress routers to destinations that do not run MPLS.

## Create MPLS Tunnels between PE routers

Before MPLS, a provider would run eBGP connections with its customers, and iBGP connections inside the provider network.\
Using iBGP requires an iBGP full mesh (administrative nightmare) or the use of route-reflectors or confederations (it may still be hard to configure in large networks and requires all routers to speak BGP)\
The more elegant solution is to use MPLS between the PE routers and to make iBGP connections only between the PE routers.

If we do not use MPLS between the PE routers, all traffic will be black-holed when it reaches the first non-BGP speaker, which will not know how to route the packet. If we use MPLS, the PE routers will encapsulate the traffic that must pass through the network and out to another AS in MPLS packets that point toward the egress PE router. This router will know how to forward the traffic using destination IP. There are some rules though:

1. iBGP connections between routers running MPLS must be brought up using the loopback addresses

   ```
   PE-RTR(config-router)# neighbor PE-ROUTER update-source LOOPBACK-ADDRESS
   ```

   If we use the interface IP to establish the neighbor relationship, PHP will kick in and it will POP the label before it reaches the last PE router, and the provider router in the MPLS cloud will not know about the BGP destination. You can also use explicit NULL advertisements to avoid this.
2. set the next-hop self for iBGP connections

   ```
   PE-RTR(config-router)# neighbor PE-ROUTER next-hop-self
   ```

   Normally, eBGP learned routes are advertised with an unchanged NEXT\_HOP value so we must first change this. Also, NEXT\_HOP value must match the egress PE router’s loopback address because this is used in LDP. Both requirements are solved using the above command.

You can also use route-maps to change the next-hop to match the LDP Router ID.

## Create the Layer 3 MPLS VPNs

### Define VRF for each customer

Using VRFs, you can have separate routing tables for each customer. Interfaces are assigned to a VRF, but by default they all belong to the global routing table.

1. Define VRF:

   ```
   PE-RTR(config)# ip vrf VRF-NAME
   ```
2. Assign interfaces to the VRF:

   ```
   R(config-if)# ip vrf forwarding VRF-NAME
   ! must re-set the IP Address on the interface
   ```
3. Verify:

   ```
   PE-RTR# show ip vrf [detail]
   PE-RTR# show ip route vrf {VRF-NAME|*}
   PE-RTR# {ping|telnet|traceroute} vrf VRF-NAME IP-ADDR
   PE-RTR# show mpls ip bindings
   ```

When you use VRFs without any MPLS or MP-BGP involved, it is called a VRF Lite. It’s simply the possibility to use separated/virtualized routing tables in a single router.

### Create VPNv4 routes and advertise them in MP-BGP

#### **RD – Route Distinguisher**

The PE routers will learn the customer routes via static routes, an IGP routing protocol or eBGP. These routes are then redistributed into MP-BGP as VPNv4 routes. A VPNv4 is created by adding a prefix to the customer route in order to make it unique in the MP-BGP table. The prefix is called RD(Route Distinguisher) and is usually in the format AS:NN or IP-ADDR:NN. The full VPNv4 prefix is RD:CUSTOMER-PREFIX/MASK-LEN. (Or AS:NN:CUSTOMER-PREFIX/MASK-LEN)\
To configure a VRF and set it up with an RD, use the following commands:

```
PE-RTR(config)# ip vrf VRF-NAME
PE-RTR(config-vrf)# rd RD
```

You will most likely need to add interfaces to the VRF:

```
PE-RTR(config-if)# ip vrf forwarding VRF-NAME
```

Assigning an interface to a VRF will remove its configured IP Address!

#### **RT – Route Target**

When routes are redistributed from the VRF into MP-BGP, they also get an extended Community attached, the Route Target Export. Similarly, on the other PE routers, routes are redistributed from MP-BGP into the customer VRF only if they have the same Route Target value as the VRF’s Route Target Import.\
A VPNv4 route can have more Route Targets assigned to it. As long as one of them matches one of the Route Target Import values on the other PE’s VRF, it will be redistributed into the customer VRF. These routes are only known by the PE routers. Inside the MPLS cloud the switching is done based on MPLS labels only

To setup the Route Targets specific for a VRF, use:

```
PE-RTR(config)# ip vrf VRF-NAME
PE-RTR(config-vrf)# route-target export EXPORT-RT
PE-RTR(config-vrf)# route-target import IMPORT-RT
! or in one single command if you use the same value:
PE-RTR(config-vrf)# route-target both EXPORT-IMPORT-RT
```

For intranet VPNs, The export RT should match the import RT on the other PEs and vice-versa. For more complex scenarios (extranet VPNs) you can use any combination of Route Targets.

#### **Connect PE routers via MP-BGP and enable VPNv4 address family**

To connect 2 or more customer sites, you need to create iBGP sessions between their respective PE routers. Otherwise, they won’t be able to exchange VPNv4 routes. To do this, use the following steps:

```
R(config)# router bgp AS-NUMBER
PE-RTR(config-router)# neighbor PE-LOOPBACK-ADDR remote-as AS-NUMBER
PE-RTR(config-router)# neighbor PE-LOOPBACK-ADDR update-source LOOPBACK-INTERFACE
```

The PE routers can connect via BGP in a Full Mesh, Hub and Spoke or any other architecture that is needed

To exchange VPNv4 routes in MP-BGP, you need to enable it in the BGP process:

```
PE-RTR(config-router)# address-family vpnv4
PE-RTR(config-router-af)# neighbor PE-LOOPBACK-ADDR activate
! activates the vpnv4 address-family for this neighbor
PE-RTR(config-router-af)# neighbor PE-LOOPBACK-ADDR send-community extended
! allows sending RT as an extended community
```

During the VPNv4 exchange, PE routers also agree on a “VPN label”. This label will be

#### **Redistribute customer routes into MP-BGP**

For VPNv4 routes to be created, you need to have IPv4 routes in the BGP process specific to the customer VRF. To do this, you must enable the address family for that specific VRF and redistribute the VRF routes into MP-BGP.

```
! On the PE router:
PE-RTR(config)# router bgp AS-NUMBER
PE-RTR(config-router)# address-family ipv4 vrf VRF-NAME
PE-RTR(config-router-af)# redistribute {static | rip | ospf PROC | eigrp AS | bgp} ... 
```

### Advertise customer routes

Since MP-BGP will advertise the customer routes from PE to PE, we need a way to advertise customer routes from CE to PE and redistribute them in MP-BGP, as well as a way to get the routes from MP-BGP and advertise them from PE to CE. You can use static routes or a dynamic IGP to achieve this.

{% tabs %}
{% tab title="Static Routes" %}
On the PE router:

```
! Customer Destination pointing to the CE Router:
PE-RTR(config)# ip route vrf VRF-NAME DESTINATION MASK NEXT-HOP-ADDR
```

Then redistribute the static or connected routes into MP-BGP:

```
PE-RTR(config)# router bgp AS-NUMBER
PE-RTR(config-router)# address-family ipv4 vrf VRF-NAME
PE-RTR(config-router-af)# redistribute static ...
PE-RTR(config-router-af)# redistribute connected ...
```

On the CE router, you can set a default gateway or more specific routes via the PE router.
{% endtab %}

{% tab title="RIP" %}
VRF aware RIP is configured inside the global RIP process, and each command can be configured for each VRF (network, version, auto-summary, etc):

```
PE-RTR(config)# router rip
PE-RTR(config-router)# address-family ipv4 vrf VRF-NAME
PE-RTR(config-router-af)# network NETWORK-ADDR
PE-RTR(config-router-af)# ...
! Redistribute routes from BGP into VRF RIP:
PE-RTR(config-router-af)# redistribute bgp...
```

Redistribute each VRF RIP into MP-BGP:

```
R(config)# router bgp AS-NUMBER
R(config-router)# address-family ipv4 vrf VRF-NAME
R(config-router-af)# redistribute rip ...
```

On the CE routers, RIP is configured normally.
{% endtab %}

{% tab title="EIGRP" %}
VRF aware EIGRP is configured inside one global EIGRP process, but each command can be configured for each VRF (network, auto-summary, etc) and the EIGRP AS-NUMBER is not inherited for each VRF, instead it must be manually configured:

```
R(config)# router eigrp GLOBAL-AS
R(config-router)# address-family ipv4 vrf VRF-NAME
R(config-router-af)# network NETWORK-ADDR WILDCARD
R(config-rotuer-af)# autonomous-system VRF-AS
R(config-router-af)# ...
! Redistribute routes from BGP into VRF EIGRP:
R(config-router-af)# redistribute bgp ... 
```

Redistribute each EIGRP VRF into MP-BGP:

```
R(config)# router bgp AS-NUMBER
R(config-router)# address-family ipv4 vrf VRF-NAME
R(config-router-af)# redistribute eigrp VRF-AS...
```

On the CE routers, EIGRP is configured normally.
{% endtab %}

{% tab title="OSPF" %}

When using OSPF, between the PE and the CE, you must use different OSPF processes for each VRF, in addition to the OSPF process used in the MPLS cloud. Also, each OSPF process must have a different Router ID.

```
R(config)# router ospf PROC vrf VRF-NAME
R(config-router)# router-id ROUTER-ID
R(config-router)# network NETWORK-ADDR WILDCARD area AREA
R(config-router)# ...
! Redistribute routes from BGP into VRF EIGRP:
R(config-router)# redistribute bgp AS-NUMBER ... 
```

You will have to redistribute each OSPF VRF into MP-BGP:

```
R(config)# router bgp AS-NUMBER
R(config-router)# address-family ipv4 vrf VRF-NAME
R(config-router-af)# redistribute ospf PROC vrf VRF-AS
```

On the CE routers, OSPF is configured normally.

**Domain ID**

The OSPF Domain ID is usually the same as the process ID, written in IP-ADDRESS format (Ex: 0.0.0.1).If the Domain ID differs between PE customer VRFs, the routes are redistributed as external, but if the domain ID is the same, OSPF internal routes from other sites are redistributed as InterArea (O IA) routes.\
You can verify the Domain ID with

```
R# show ip ospf [PROC-ID] | i Domain
   Domain ID type 0x0005, value 0.0.0.1
```

To change it, use:

```
R(config)# router ospf PROC-ID vrf VRF-NAME
R(config-router)# domain-id DOMAIN-ID
```

**Sham Links**

When running OSPF as the customer protocol, the following scenario may have unwanted results. Since the customer routes redistributed in OSPF from BGP appear as external or InterArea routes (as seen before), if the 2 sites have a backdoor connection that runs OSPF, then it is highly probable that we will learn the same routes over this backdoor connection as IntraArea routes (O). As you know, OSPF will always prefer IntraArea routes to InterArea and External Routes, so there is no way you can force the customer traffic to use the VPN rather than the backdoor (usually used for backup) route. Unless you use sham-links! Sham links are virtual links that are used to connect 2 customer PEs. By learning the routes over the sham link, they will be considered IntraArea (O) and may be preferred based on the cost to the routes learned via the backdoor links.\
To configure a sham link, we will need one additional /32 loopback on each PE. These loopbacks will not be advertised using OSPF into the customer VRF, instead they will be advertised into BGP. On the PEs we will configure:

```
PE1(config)# interface LO1
PE1(config-if)# ip address IP-ADDR1 255.255.255.25
PE1(config)# router bgp AS-NUMBER
PE1(config-router)# network NETWORK1 mask 255.255.255.255
PE1(config)# router ospf PROC-ID1 vrf VRF-NAME1
PE1(config-router)# area AREA-ID1 sham-link IP-ADDR1 IP-ADDR2 [cost COST]
! IP-ADDR1 is configured on PE1, and IP-ADDR2 si configured on PE2
```

and

```
PE2(config)# interface LO1
PE2(config-if)# ip address IP-ADDR2 255.255.255.25
PE2(config)# router bgp AS-NUMBER
PE2(config-router)# network NETWORK2 mask 255.255.255.255
PE2(config)# router ospf PROC-ID2 vrf VRF-NAME2
PE2(config-router)# area AREA-ID1 sham-link IP-ADDR2 IP-ADDR1 [cost COST]
```

You can verify the sham-link status using:

```
R#show ip ospf sham-links
```

Now if you look at the OSPF routes on the CE the InterArea (O IA) routes should be replaced by the IntraArea (O) routes. Of course, the metric of the routes received via the sham links should be better than via the backdoor link.

{% endtab %}
{% endtabs %}

### **Verify**

```
R# show ip bgp vpnv4 [VRF-NAME | all]
R# show ip bgp vpnv4 rd RD labels
R# show mpls ip binding
```

### How a packet moves in an MPLS VPN network

When the PE routers become peers, and exchange VPNv4 rotues, they also agree on a “VPN label” (or “BGP Label”)to be used for each prefix. This label will be at the bottom of the stack (BoS:=1) and will be attached to all packets that are intended for a specific prefix inside the customer network.\
At the same time, since the PE routers had their “next-hop self” configured, the VPNv4 routes received by each PE have their next hop set to the other PE’s IP address. This address is advertised inside the provider cloud and via LDP, son when adding the MPLS header to a packet coming/going to the customer network, the router must add 2 labels:

* Transport/IGP Label – as the label used to forward traffic towards the destination PE router
* VPN/BGP Label (BoS=1) – identifies the customer VPN when reaching the PE router

The Transport label is swaped as the packet goes through each LSR in the LSP. Usually, just before reaching the PE router, the Transport label is popped (due to PHP). Now, if a PE router has multiple VRF customers, how would it know where to forward the packet? IP information is not good enough, because address spaces can overlap. That’s why we used VPNv4. This is where the VPN label comes into play. It will help the router understand what VRF it is destined for. Since the label is agreed during VPNv4 route exchange, the label is bind to a specific VRF.\
Once this information is clear, the router can strip the MPLS labels and forward the packet based on its IP address.


# Multicast 101

## Multicast Addressing

### Layer 3 Addressing

Multicast IP packets are sent from the source to a destination IP in the range 224.0.0.0/4 (224.0.0.0 – 239.255.255.255). It means every multicast IP address starts with 1110 bits. This range is further subdivided in several blocks. The most important ones are:

* **Local Network Control Block (aka Link Local) – 224.0.0.0/24** (224.0.0.0 – 224.0.0.255). Packets with an address in this range are local in scope are not forwarded by routers. They are usually sent out with a TTL of 1.
* **Internetwork Control Block – 224.0.1.0/24** (224.0.1.0 – 224.0.1.255)
* **Source Specific Multicast – 232.0.0.0/8** (232.0.0.0 – 232.255.255.255)
* **GLOP – 233.A.S.0/24** (233.A.S.0 – 233.A.S.255), where A and S are each 8 bits of the 16 bits AS Number. These addresses are used by companies who have  already been allocated an AS Number and can use it for multicasting over Internet
* **Administratively Scoped Addresses – 239.0.0.0/8** (239.0.0.0 – 239.255.255.255) – These addresses are considered local to the site and should be considered private traffic (similar to RFC1918 IP Addresses)

For a complete list, see [IANA](https://www.iana.org/assignments/multicast-addresses/multicast-addresses.xml)

### Layer 2 Addressing

#### **Ethernet**

Of course, each layer 3 packet is encapsulated in a Layer 2 frame. On Ethernet, the frame has the MAC address of the sender as the source and a special address derived from the Layer 3 destination IP as the destination.

The Layer 2 Multicast Addresses has an OUI of **01-00-5E, followed by a 0(zero)  and 23 more bits**. The last 23 bits will be the same as the the last 23 bits in the Layer 3 Multicast Address. But the multicast address is always in the format 1110 + 28 more bits. That means that 5 bits are lost in the process of conversion. This also means that 2^5=32 Layer 3 addresses have the same Layer 2 address. They are:

```
224.X    .Y.Z
224.X+128.Y.Z
225.X    .Y.Z
225.X+128.Y.Z
...
239.X    .Y.Z
239.X+128.Y.Z
```

#### **Frame Relay**

On Frame Relay networks several issues might appear when routing multicast traffic. First of all, you need to set the **broadcast** keyword on the frame-relay maps where you want to send multicast frames. Most of the problems can be solved using point-to-point subinterfaces, but if this is not the case, then other issues may arise when using hub-and-spoke topology with multipoint interfaces.

* When sending multicast traffic from one spoke to another, the same interface on the hub should be both the incoming interface and part of the OIL, which is impossible. The solution is to set the pim nbma-mode, where the JOINs are tracked per DLCI, and not per interface. This will allow traffic to enter one DLCI and exit another DLCI on the same interface. Since the PIM JOINs are tracked, only PIM-SM is supported for this configuration:

  ```
  R(config-if)# ip pim nbma-mode
  ```
* Also, since the spokes will not see each other, there is no way for one spoke to do a PRUNE OVERRIDE if it needs traffic. This means that the multipoint interface will be pruned and multicast traffic will not reach the spoke that needs it

## Protocols used

Multicast uses 2 types of protocols. First, there are the protocols used by the receivers to signal to the routers that they want to receive traffic for a certain group – [IGMP](https://nyquist.eu/igmp-101/). Then, there is the protocol used by the routers to build distribution trees that connect the sources to the receivers – PIM, DVMRP (Distance Vector Multicast Protocol) or MSDP (Multicast Source Discovery Protocol) belong to this category.

## Multicast Routing Table

### Reverse Path Forwarding

Reverse Path Forwarding is the primary loop prevention mechanism used in multicast routing. It is a simple mechanism but it successfully accomplishes the task of not duplicating packets in a multipath network.

Unicast routing is normally done by forwarding the packet based on its Layer 3 destination address. In contrast, Multicast routing forwards the packet based on its Layer 3 source address, by forwarding the packet away from the source.

Reverse Path Forwarding kicks in when a multicast packet is received. The router will run an RPF check on each packet to determine if the interface it came in was the interface that points to the source of the traffic. If it does, then it forwards the packet out all\* other multicast interfaces, except the incoming interface. If the RPF check fails, the router just drops the packet. It means it didn’t come from the source, and by forwarding it, we would create a loop.\
You can see how a RPF check is resolved issuing the command:

```
R# show ip rpf IP-ADDRESS
! if it says failed, traffic is not forwarded
```

Additional RPF settings can be configured with commands that start with:

```
R(config)# ip multicast rpf ...
```

### mRoute Table

A router can only have 1 incoming interface for any entry in the mroute table. If the multicast traffic comes from multiple interfaces, the route with the best cost is chosen \[AD/metric]. In case of equal cost load balancing, the entry with the highest next-hop IP is used as incoming interface. All other interfaces will fail the RPF check.\
Different multicast protocols use different methods to perform this RPF check. PIM (Protocol Independent Multicast) uses the unicast routing table to determine the interface that points to the source. Other protocols can maintain a separate Multicast Routing table.

The multicast routing table consists of different (\*,G) and (S,G) entries. Each entry has an incoming interface and an OIL (Outgoing Interface List). The Incoming interface cannot be at the same time in the OIL, too. If a packet passes the RPF check it is sent out all interfaces in the OIL.

To enable multicast routing on a Cisco router, use:

```
R(config)# ip multicast-routing
```

To see the mroute table, use:

```
R# show ip mroute
```

The entries in the routing table have a list of flags associated to them. This list

```
D - Dense = group works in dense mode
S - Sparse - group works in sparse mode
C - Connected  = there is a directly connected receiver
L - Local = the router itself is a receiver
P - Pruned = a PRUNE was sent upstream (no neighbors or receivers downstream)
T - Shortest Path Tree = the group uses a SPT. 
        PIM-DM - Always on
        PIM-SM - Only on (S,G) entries
J - JOIN = 
        PIM-DM - can be ignored
        PIM-SM - (*,G) signals a SPT switchover. The next packet will be routed using the (S,G). Recalculate threshod once a second
               - (S,G) signals a SPT switchover. Recalculates threshold once a minute.
F - REGISTER = there's a source on this interface
R - information in the (S,G) applies to the shared tree
```

### Static mRouting

When a RPF check fails, the incoming packet is dropped. The RPF check is done by finding the outgoing interface of the best route towards the source in the unicast routing table. When this is not the intended interface we have 2 options:

1. Manipulating the unicast table – usually not a good idea
2. Manually setting the incoming interface:

   ```
   R(config)# ip mroute SOURCE MASK { PREVIOUS-HOP | INCOMING-INTERFACE} [AD]
   ```

   You can configure floating static mroutes if you specify an AD higher than the route in the routing table.

   This entry will be used in the mroute table. It is not an actual static route, but rather a static change of the RFP result.

## Multicast Filtering

### Multicast Boundary

Multicast traffic can be filtered on an interface using:

```
R(config-if)# ip multicast boundary ACL [filter-autorp] {in|out}
! ACL - permitted groups are allowed
! filter-autorp - also filter auto-rp groups (224.0.1.39, 224.0.1.40)
```

Both data and control plane traffic are filtered with this command and it works in both directions.

### TTL Scoping

A less flexible option for filtering multicast packets is the TTL scoping feature. In this setup, you can configure a TTL value on the interface beyond which any multicast packet that reaches that interface is dropped if the TTL value of the packet is less than configure on the interface. Normally packets are dropped when the TTL value reaches 0, but this way you can drop packets with a higher TTL as well. This only applies to multicast packets, but it applies to all of them. To enable TTL scoping on an interface, use:

```
R(config-if)#ip multicast ttl-threshold [TTL-VALUE]
! Default:0
```

## Multicast Rate Limiting

Multicast traffic can be rate-limited on each interface:

```
R(config-if)# ip multicast rate-limit {in|out}  [video|whiteboard]  [group-list ACL-GRP] [ source-list ACL-SRC] KBPS
! video, whiteboard - limits traffic only for known UDP ports
! without group-list and source-list: all multicast traffic si rate-limited
! group-list - only limits traffic for groups permitted by ACL
! source-list - only limits traffic from sources permitted by ACL
```

## Mapping multicast and broadcast

Some applications only support sending or receiving broadcast traffic. Whatever you do, broadcast traffic will not be forwarded by a router. What you can do is to convert the broadcast traffic into unicast or multicast.

### Convert broadcast to unicast – Helper Address

To convert it to unicast, use the following command on the incoming interface:

```
R(config-if)#ip helper-address DESTINATION-ADDR
```

Broadcast traffic will be forwarded as unicast to the DESTINATION-ADDR.

### Forwarding UDP Broadcasts

But what traffic? By default, a router will forward only broadcasts for these UDP ports:

* TFTP – UDP 69
* DNS – UDP 53
* Time service – UDP 37 – This is not NTP
* NetBIOS Name Server – UDP 137
* NetBIOS Datagram Server – UDP 138
* BOOTP – UDP 67,68
* TACACS – UDP 49
* IEN-116 Name Service – UDP 42

You can add additional services to this list with the global command:

```
R(config)# ip forward-protocol udp PORT
! Additional protocols besides UDP can also be forwarded
```

### Convert broadcast to multicast – Helper Map

To convert broadcast to multicast and the other way around, use the **multicast helper-map** command.\
On the router that converts Broadcat to Multicast:

```
R(config-if)# ip multicast helper-map broadcast GROUP-ADDR EXTENDED-ACL [ttl TTL]
! Use EXTENDED-ACL to limit the traffic type to certain UDP Ports
```

On the router that converts Multicast to Broadcast:

```
R(config-if)# ip multicast helper-map GROUP-ADDR BROADCAST-ADDR EXTENDED-ACL [ttl TTL]
! Use EXTENDED-ACL to match only specific UDP traffic
```

### Directed Broadcast

Directed broadcasts are IP packets that are destined to the broadcast address of a remote subnet. They have the host part of the destination set to all 1s and they are routed just like unicast packets until they reach the router that has that subnet directly connected. By default, this router will drop the packet and will not forward it on the subnet. You can forward the packets as 255.255.255.255 broadcasts if you enable the following command on the interface connected to the end subnet:

```
R(config)# ip directed-broadcast [ACL]
! If ACL is used, unmatched traffic is dropped
```

### Broadcast Address

Normally, the broadcast address used by a router is 255.255.255.255. Older implementations use the 0.0.0.0 address as the broadcast address and if you want to switch the broadcast address globally you can modify the config-register and reload the router.\
However, if you want to specify a different broadcast address to be used only on some interfaces, you can use the command:

```
R(config-if)# ip broadcast-address BROADCAST-ADDR
```


# PIM 101

## PIM basics

PIM is a protocol used to create loop-free multicast delivery trees from the source to the receivers. There are 2 different modes for PIM – Dense Mode and Sparse Mode, and a few other variations: PIM Sparse-Dense Mode, Bidirectional PIM and Source Specific PIM.\
PIM uses the unicast routing table to do the [RPF](/multicast/multicast-101#reverse-path-forwarding) check of incoming multicast packets.\
PIM will not forward multicast traffic out the interfaces where PIM is not enabled. By default, no interfaces are running PIM until a PIM mode is defined on them:

```
R(config-if)# ip pim {dense-mode|sparse-mode|sparse-dense-mode}
```

To see a list of interfaces running PIM,use:

```
R# show ip pim interfaces [INTERFACE][detail|stats|count]
```

### PIM Neighbors

PIM uses a neighbor discovery mechanism by sending HELLOs every 30 sec (default) to 224.0.0.13. The interval can be changed with

```
R(config-if)# ip pim query-interval HELLO-TIME [msec]
!Default: 30 sec
```

PIM uses neighbor adjacencies to build the distribution tree. By default, PIM will accept any PIM-enabled router as a neighbor, but you can prevent that with:

```
R(config-if)# ip pim neighbor-filter ACL
```

To see a list of PIM neighbors, use:

```
R# show ip pim neighbors
```

### PIM DR

A PIM DR is elected for both PIM Dense Mode and PIM Sparse Mode. While on PIM Dense Mode it has no explicit role (other than being the IGMPv1 Querier), in PIM Sparse Mode, the DR will Register a source to the Randezvous Point.\
HELLOs are also used to elect the PIM DR. The DR is elected on a multiaccess network as the router with the highest priority and then the highest IP Address. To set the priority, use:

```
R(config-if)# ip pim dr-priority PRI
!Default Priority = 1. Highest wins
```

## PIM Dense Mode

### Building the Shortest Path Tree

PIM Dense Mode was the initial method of delivery of multicast traffic. It uses a “Flood and Prune” mechanism. This means that by default, PIM DM will send traffic out all PIM-enabled interfaces as long as there is a PIM Neighbor or an IGMP client that joined the destination multicast group. By using this approach, the network is “flooded” with multicast traffic until it reaches the intended clients. A multicast “Source Distribution Tree” a.k.a “Shortest Path Tree” is built from the source as the root. More exactly, whenever multicast traffic arrives at the router and it passes the [RPF ](/multicast/multicast-101#reverse-path-forwarding)check, it will populate it’s mroute table with an (S,G) entry. The interface where multicast traffic arrived will be the incoming interface, and all other interfaces where PIM neighbors exist or where directly connected clients exist (known via IGMP) will be added to the OIL (Outgoing Interface List).

To prevent unwanted multicast traffic, a Router can send a PRUNE message upstream, telling its upstream PIM neighbor to stop sending the traffic this way. A router will send PRUNE messages upstream when:

* traffic is arriving on a non-RPF poin-to-point interface
* it is a leaf router (no PIM neighbors downstream) and no directly connected receivers
* it is a non-leaf router on a point-to-point link that has received a PRUNE message from the neighbor
* it is a non-leaf router on a LAN segment that has received a PRUNE from a neighbor an no one overrides that PRUNE.

When a PRUNE is received, a 3 seconds (default) prune delay is started. When a PRUNE is sent on a LAN, if somebody still wants the traffic it will send a PRUNE OVERRIDE message. If such a message is received before the prune delay times out, the traffic is not pruned.

A PRUNE times out every 3 minutes by default. When it times out, the PIM router will start to send multicast traffic again on the interface. A new PRUNE message is needed for the router to stop the flooding.\
To prevent the periodic flooding, PIM implemented a PIM STATE REFRESH message that is sent downstream from the routers closest to the source. If traffic is needed on an interface, the downstream routers can send a GRAFT message upstream to request traffic. This message is acknowledged with a GRAFT-ACK before traffic is forwarded on that interface.

To enable PIM Dense Mode, use the following command:

```
R(config-if)# ip pim dense-mode
! It will also enable IGMP
! If multicast-routing is not enabled, it gives a warning
```

It is very important to understand the data in the following show command:

```
R# show ip mroute MULTICAST-GROUP
! (*,G) = Container entry, not used for forwarding in PIM DM
(*, MULTICAST-GROUP), 00:22:41/stopped, RP 0.0.0.0, flags: D
  Incoming interface: Null, RPF nbr 0.0.0.0
  Outgoing interface list:
    INTERFACE1, Forward/Dense, 00:22:41/stopped
    INTERFACE2, Forward/Dense, 00:22:41/stopped
! (S,G) entry is used for forwarding:
(MULTICAST-SOURCE, MULTICAST-GROUP), 00:00:47/00:02:12, flags: PT
  Incoming interface: INTERFACE1, RPF nbr UPSTREAM-NEIGHBOR
  Outgoing interface list:
    INTERFACE2, Prune/Dense, 00:00:47/00:02:12
```

The first entry is always a (\*,G) group, but this is not actually used for forwarding. Instead, it is just a Group that lists all interfaces that could be part of the OIL. Forwarding is don based on the (S,G). If there are multiple sources that send to the same multicast group, then we’ll have multiple (S,G) entries. The information in the (S,G) entry tells us that traffic was received on INTERFACE1 and was forwarded out INTERFACE2, but we recently received a PRUNE message on that interface. No clients on this branch of the tree. If there where no PIM neighbors either, then the OIL would have said **“Null”**. If we had a client on this branch than it would’ve said **“Forward/Dense”**. If the traffic on that interface is pruned then, just as in our example, it says **“Prune/Dense**.\
If the incoming interface in the (S,G) entry is **“Null”**, it means that the traffic from this source failed the [RPF](/multicast/multicast-101#reverse-path-forwarding) check. This is usually due to the fact that our routing table points to another interface to reach this source. You can make the traffic pass the RPF check if you modify the routing table accordingly, or if you force the router to accept the traffic on this interface, using a static mroute:

```
R(config)# ip mroute SOURCE MASK {PREVIOUS-HOP|INCOMING-INTERFACE}
! PREVIOUS-HOP in multicast equals to NEXT-HOP in unicast routing table
```

If there are multiple sources that transmit, you can see a summary of the mroute table with:

```
R# show ip mroute summary
(*, MULTICAST-GROUP), 01:14:12/stopped, RP 0.0.0.0, OIF count: 3, flags: D
  (SOURCE1, MULTICAST-GROUP), 00:00:46/00:02:13, OIF count: 2, flags: T
  (SOURCE2, MULTICAST-GROUP), 00:06:47/00:01:02, OIF count: 2, flags: T
```

### PIM Forwarder

When sending traffic downstream, PIM ASSERT messages are used to determine the PIM forwarder. If a router receives a multicast packet on an interface that is in the OIL (Outgoing Interface List), it means there is another router forwarding traffic on the same segment. This will result in duplicate packets being sent to the receivers. To solve this, the routers will send a PIM ASSERT message on the LAN. This message contains the metric information \[AD/metric] to the source of the traffic. The router with the best metric will win. In case of a tie, the router with the highest IP address will win.\
The only way you can change the PIM Forwarder is by changing the routing table so it uses a better \[AD/metric] in the ASSERT messages.

## PIM Sparse Mode

PIM Sparse mode is defined in [RFC 4601](https://tools.ietf.org/html/rfc4601). It uses an explicit JOIN mechanism where no traffic is sent out an interface, unless an explicit JOIN was received on that interface.\
To enable PIM Sparse Mode on an interface, use:

```
R(config-if)# ip pim sparse-mode
```

Sparse Mode considers that no interface needs the multicast traffic at the beginning, and each router that has directly connected receivers must explicitly send a JOIN upstream. Interfaces that receive JOINs are added to the Forwarding list in the OIL, but they timeout after 3 minutes. JOINs are usually sent every minute to keep the interface forwarding.

Since the receivers must explicitly JOIN the distribution tree, there is a need for a Rendezvous Point that will keep track of all receivers and sources in the network.

PIM-SM is basically split into 2 forwarding trees. One from the source to the RP and another from the RP to the receivers.

Another difference between PIM Dense Mode and PIM Sparse Mode is that the entries in the mroute table are generated by multicast traffic in PIM DM, and by PIM JOINs in PIM SM.

### Building the Shared Tree

The router that receives IGMP JOINs (PIM DR if more than one) sends a PIM JOIN upstream to the RP. All routers in the path insert (\*,G) entries and send the PIM JOIN upstream until it reaches the RP. The (\*,G) entries will have an incoming interface set to the interface towards the RP and an OIL containing the interface where the JOIN was received.

If traffic is no longer needed, a (\*,G) PRUNE must be sent upstream. When a PRUNE is received, the interface is excluded from the OIL. If the OIL is empty (NULL) then a PRUNE is sent upstream. A (\*,G) PRUNE will also prune all (S,G) entries

### Source Registration and Building the Source Distribution Tree

Since the RP is the root of the distribution tree, there must be a way for the sources to send the traffic to the RP for distribution. This is done via a registration process. When a source starts to transmit, the DR on its segment encapsulates the packets in a PIM REGISTER message, which is unicasted to the RP.\
When the RP receives the message from the PIM DR, it will acknowledge it with a REGISTER STOP message and will insert an (S,G) entry into the mroute table. Now, only the DR and the RP know about the (S,G) entry.\
The RP can be configured with a list of allowed sources that can register to it:

```
R(config)# ip pim accept-register ACL
! The ACL shoud be extended. It matches the SOURCE and the GROUP
```

Once the RP has a (\*,G) entry in its mroute table (meaning there is a receiver a.k.a the shared tree was built), it sends its own PIM JOIN message upstream to the source, making all routers on the upstream path to add (\*,G) entries with OIL pointing to the RP.

By default, the PIM DR sources the register message from the interface closest to the RP. You can change this with:

```
R(config)# ip pim register source INTERFACE
```

### SPT Switchover

A router that has directly connected receivers can switch from the Shared Tree to the SPT to the source when the traffic on the shared tree exceeds a certain threshold:

```
R(config)# ip pim spt-threshold {KBPS |infinty}
! Default: 0 = always
```

By default, this threshold is 0 for Cisco Routers, meaning that it will always switchover. It could be set to never switchover by using the *infinity* keyword.\
When this router decides to switchover, it sends a (S,G) JOIN upstream. It is possible, at some point in the shared tree, that the SPT tree and the Shared Tree will diverge. Routers will send a (S,G) PRUNE up the shared tree when the upstream neighbor for (S,G) is different than the RP. This will prune the shared tree. An (S,G) PRUNE will only prune the (S,G) entry and will not generate a PRUNE upstream.

### PIM-SM mRoute Table

Unlike PIM-DM, in PIM-SM (\*,G) entries are used to forward multicast traffic down the shared tree. (\*,G) entries are created when JOINs are received. The incoming interface of a (\*,G) entry is set to the RPF interface that points to the RP, and the OIL contains interfaces where PIM JOINs for this group were received.\
(S,G) entries are created as a result of explicit JOINs for (S,G). (S,G) JOIN messages are sent by a router that is doing SPT Switchover or by the RP when a REGISTER message is received.

In PIM-SM pruned interfaces are removed from OIL so they will only show up in the OIL if traffic needs to be sent on them. Therefor they always appear as **Forward/Sparse**

### Setting the Rendezvous Point (RP)

PIM Sparse Mode uses a shared tree to deliver traffic to the receivers. Without a RP, PIM SM will not work.\
The RP can be set in 3 ways: manual, with Auto-RP, or with BSR.

To prevent unwanted RPs in the network, a router can be configured to ignore JOIN messages addressed to other RP then the ones specified in:

```
R(config)# ip pim accept-rp {RP-ADDR [ACL] | auto-rp}
```

This means that the router will ignore these JOIN messages so they will not generate entries in the mroute table.

#### **Manual**

```
R(config)# ip pim rp-address RP-ADDR [ACL] [override]
```

When setting the RP manual, it can be done for all groups or for a subset (matching the ACL). Normally, if the RP is set in any other way, it will be considered a better choice than the manual setting. Use the *override* keyword to use the statically defined RP even if another one is set using Auto-RP or BSR.

#### **Auto-RP**

Auto RP is a feature implemented first by Cisco. It helps routers find the RP in a network. There are 2 components in the design: the RP candidates and the mapping agent.\
First, the RP candidates must announce themselves at the 224.0.1.39 address.

```
R(config)# ip pim send-rp-announce INTERFACE scope TTL [group-list ACL][interval SEC]
!The INTERFACE must run PIM (including Loopback)
```

Then the mapping agents must be configured to join the 224.0.1.39 address and listen to the announcement. They will announce the mappings to the 224.0.1.40 group, an address that all Cisco PIM-SM routers join by default. To configure a mapping agent, use:

```
R(config)# ip pim send-rp-discovery INTERFACE scope TTL [group-list ACL][interval SEC]
```

If there are multiple RP candidates, the Mapping Agent will choose the router with the highest IP Address.\
The mapping agents can filter incoming RP candidates announcements, using:

```
R(config)# ip pim rp-announce-filter [rp-list RP-ACL] [group-list GROUP-ACL]
```

The GROUP-ACL must match the ACL used by the RP Candidate in its announcement, otherwise it will be completely filtered.\
The mapping agent uses Dense Mode to send RP information, and we need RP to join any groups. This ends up in a circular logic. The available solutions are:

* Run PIM in sparse-dense mode
* Statically define the RP for the Auto-RP Groups. (224.0.1.39 and 224.0.1.40) – this defeats the purpose of auto-configuration
* Set the router to be an auto-rp listener, which is a workaround that makes the router work in dense mode just for those 2 addresses, and in sparse mode for all other multicast groups.

  ```
  R(config)# ip pim rp-autolistener
  ```

To see the RP mappings, use:

```
R# show ip pim rp [mapping]
```

You can filter auto-rp messages going out an interface with:

```
R(config-if)# ip multicast boundary filter-autorp [ACL]
```

#### **BSR – BootStrap Router**

BSR is an open standard defined in PIMv2 RFC, used to automate the process of electing a RP.\
Similar to Auto-RP, there are several candidate RPs and one or more BSR routers that will advertise RP to group mappings. The BSR is elected first based on the highest priority and then by the highest IP address. To define a router as a BSR Candidate, use:

```
R(config)# ip pim bsr-candidate INTERFACE [HASH-LEN] [PRI]
! Default PRI=0 - Highest wins. If tied, the higher IP wins.
```

The BSR algorithm tries to distribute groups more or less equal between all Candidate RPs, but at the same time to have all routers select the same RP for a group. For each Group Address, the BSR algorithm will select the Candidate RP with the lowest Priority (as advertised by the Candidate RP to the BSR). Since this is 0 by default, it often happens that there are multiple RP Candidates with the same lowest Priority.\
The Hash function, which can be found in [RFC 2362](https://www.ietf.org/rfc/rfc2362.txt), will generate a Hash value for each Candidate RP that is available for a Group and that has the same lowest Priority. The Candidate RP with the highest hash value will be selected as the RP. To see the result of the Hash value, use:

```
R# show ip rp-hash GROUP-ADDR
```

HASH-LEN represents the length of a mask that is ANDed with the Group Address in order to obtain a Hash value. This makes the hash function to actually use only the first HASH-LEN bits of the Group Address. Therefor, all groups that have the same HASH-LEN bits will select the same RP. By default, HASH-LEN is zero, so all groups will select the same RP (You can see that the hash value is the same for all RPs, regardless of the Group Address). If you define a HASH-LEN value, then you can have smaller continuos chunks of the multicast address space assigned to the same RP.\
For example, when using a HASH-LEN of 31, every 2 Group Address will be assigned to the same RP. For 30, every 4 Group Addresses, and so on.\
At this point you would think that using a HASH-LEN of 1 would assign half of the Groups to an RP and the other will be assigned to an RP, and the other half to another RP. But that’s not the case. Since the Multicast Address all share the same 4 bits (224.0.0.0/4) a HASH-LEN of 4 or less will result in the same result of the hash function. Therefor, the minimum HASH-LEN should be 5 if you attempt to assign half of the Group Addresses to an RP and the other half to another RP. However, the Hash results are pseudo-random and you have better chances of equal distribution as the number of different hash results grows. That is a 32 HASH-LEN will compute different values for each address and assign the group to one of the Candidate RPs.

BSR information is passed on a hop-by-hop basis to 224.0.0.13 (All PIM Routers). This is how routers will find information about the RP mappings.

RP canidates send unicast information to the BSR, because they will know the BSR address:

```
R(config)# ip pim rp-candidate INTERFACE [group-list ACL] [interval SEC] [priority PRI]
! Default PRI=0. Lowest wins!
! Default SEC=60
```

The ACL used to filter the groups it advertises for, can only have permit entries.

You can filter BSR messages going out an interface with:

```
R(config-if)# ip pim bsr-border
```

## PIM Sparse-Dense Mode

PIM Sparse-Dense Mode uses Sparse Mode when there is a RP for the group, and Dense Mode when there is no RP. It can be used to make Auto-RP work.\
To enable sparse-dense mode on an interface, use:

```
R(config-if)# ip pim sparse-dense-mode
```

There is a problem when the RP becomes unavailable. The groups will fallback to Dense Mode and the information will be flooded throughout the multicast network. To prevent this, you can:

* disable dense-mode fallback:

  ```
  R(config)# no ip pim dm-fallback
  ```
* use Auto RP listener:

  ```
  R(config)# ip pim autorp listener
  ```
* or set a “sink RP” as a last resort – a static RP that will prevent Dense Mode fallback.

  ```
  R(config)# ip pim rp-address RP-ADDR [ACL]
  ```

## Bidirectional PIM

Bidirectional PIM is based on PIM-SM but differs in the way the source sends its data to the RP. In normal PIM-SM, traffic would only flow downstream from the source to the RP on the SPT and from the RP to the receivers on the shared tree. In Bidir PIM, traffic flows only on a shared tree. Upstream from the source to the RP and downstream from the RP to the receivers.

On every network segment, a Designated Forwarder (DF) is elected to forward the traffic upstream towards the RP. There will be one DF for each RP in the network. There is no source registration in Bidir Mode. Instead, all traffic will reach the RP where it will be dropped if there are no clients.

Routers will only use (\*,G) entries in the mroute table and when performing RPF check, the interface towards the RP is used. When sending traffic towards the RP, only the DF will forward packets, all other routers will discard them.\
Bidirectional PIM will skip the RPF check, but all routers in a domain must be configured to support this feature, otherwise loops can occur.

To configure a router to support PIM, use:

```
R(config)# ip pim bidir-enable
```

Of course, the routers will need to work in Sparse or Sparse-Dense Mode, and they will need to know the RP for this traffic.\
To enable a RP for Bidirectional PIM you can use any method, just specify that it works in bidir mode:

```
! Manual:
R(config)# ip pim rp-address RP-ADDR [ACL] bidir
! Auto-RP
R(config)# ip pim send-rp-announce INTERFACE scope TTL [group-list ACL] bidir
! BSR
R(config)# ip pim rp-candidate [group-list ACL] bidir
```

## PIM SSM – Source Specific Multicast

SSM is a delivery model in which traffic is forwarded to the receivers only from those sources that they explicitly joined. Only IGMPv3 supports this kind of Membership Reports where the receiver specifies the source it wants to receive from. In PIM SSM, the routers only create (S,G) entries, but this applies only to the SSM address range. All other groups are treated as normal Sparse Mode.

```
R(config)# ip pim ssm {default | range ACL}
```

The default range is specified by IANA as 232.0.0.0/8 but any other multicast range can be specified using the ACL.\
Remember to set the IGMP version to 3 on the interfaces where receivers exist:

```
R(config-if)# ip igmp version 3
```

In SSM, there is no need for RP because the clients know the source of the traffic. The Clients generate IGMPv3 (S,G) JOINs, and the routers build the tree up to the source using (S,G) PIM JOINs.


# IGMP 101

To make multicast work, we use 2 types of protocols: one that is used to signal the receivers to the routers, and the other used by the routers to build distribution trees from the source to the receivers. In the first category we have IGMP (Internet Group Management Protocol), while in the second, PIM (Protocol Independent Multicast), DVMRP (Distance Vector Multicast Routing Protocol) or MSDP (Multicast Source Discovery Protocol).

## Versions

IGMP is used between hosts on a LAN and the routers to track the multicast receivers on that LAN. On Cisco routers, IGMP is enabled when PIM is enabled on the interface.

```
R# show ip igmp interface
```

There are 3 version of IGMP, and by default, Cisco IOS runs IGMPv2

```
R(config)# ip igmp version {1|2|3}
!Default: 2
```

Version 1 is defined in [RFC 1112](https://www.ietf.org/rfc/rfc1112.txt), version 2 in [RFC 2236](https://www.ietf.org/rfc/rfc2236.txt), version 3 in [RFC 3376](https://www.ietf.org/rfc/rfc3376.txt).

## How hosts JOIN and LEAVE a group

### IGMPv1

On a LAN where there are more than one routers, only one of them must act as the IGMP Querier and talk to the hosts on the network. In version 1, IGMP relies on PIM to select a Querier – the PIM DR will act as the Querier (highest IP). This differs from IGMPv2, where there is specific mechanism to elect the IGMP Querier (lowest IP).

#### **Joining a Group**

The IGMP Querier sends an IGMPv1 Membership QUERY to 224.0.0.1 (All Hosts). Each host that receives the Query stars a random timer (max 10 sec) and when the first one expires, that host sends an IGMP MEMBERSHIP REPORT to the Group Address that it wishes to join. Each host in the group receives this REPORT and will suppress it’s own report because the router does’n need to know who exactly needs the traffic, just if there is anyone who needs it.

```
R(config)# ip igmp query-interval INTERVAL
!how often will the router send IGMP QUERIES on the LAN. Default: 1 min
```

A host can also send unsolicited MEMBERSHIP REPORTS when they want to join a group.

#### **Leaving a Group**

In IGMPv1, leaving a group is only achieved via timeout. The Timeout is 3x Query Time, that is 3 minutes by default.

### IGMP v2

The IGMP Querier is elected in version 2 as the router with the lowest IP address. The time to wait before another router can become the acting querier can be defined with:

```
R(config)# ip igmp querier-timeout TIMEOUT
!Default: 255 sec
!how long before detecting that the acting Querier is gone
```

When IGMPv2 hosts detect IGMPv1 routers, they must fallback to using IGMPv1 and stop sending IGMPv2 messages until no IGMPv1 messages are received. When IGMPv2 routers detect IGMPv1 hosts, they will ignore GROUP-SPECIFIC-QUERIES and LEAVE messages because these are not understood by IGMPv1 hosts.

#### **Joining a Group**

The IGMP Querier sends 2 types of  IGMP Membership QUERIES in version 2. There is the GENERAL-QUERY that acts just like a v1 QUERY, but there is also a GROUP-SPECIFIC-QUERY, used to find if a specific group still has members.

Additionally, IGMPv2 GROUP-SPECIFIC-QUERIES offer the possibility to set the MAX-RESPONSE-TIME, the random timer used by the hosts before responding.

```
R(config)# ip igmp query-max-response-time MAX-TIME
!Default MAX-TIME: 1 sec => Quicker than GENERAL QUERY
```

#### **Leaving a Group**

IGMPv2 adds a LEAVE group message that is usually sent by any host when it leaves a group. When the router receives a LEAVE message, it sends a GROUP SPECIFIC QUERY to see if there are any receivers left on the LAN. If there are hosts, they will respond with a REPORT. If there is no response, the MAX REPONSE TIME will timeout (1 sec) and the router will send a new Group specific QUERY (default: 2 times). If again no response in time, the router will not forward traffic for this group on that segment. To change the defaults:

```
R(config-if)# ip igmp last-member-query-count COUNT
!Default: 2
R(config-if)# ip igmp last-member-query-interval INTERVAL
!Default: 1 second
```

You can also set the router to immediately prune the traffic. The receivers must then send JOIN messages if they still want to receive the traffic:

```
R(config)# ip igmp immediate-leave
```

### IGMP v3

#### **Joining a Group**

In IGMPv1 and IGMPv2, receivers would only join a destination group, and traffic from any source sent to this destination was going to reach the receivers. Version 3 introduces support for Source Specific Multicast (SSM), which means that receivers will only get the traffic sent by a specific source. This means that when receiving IGMPv3 JOINS, a router can create specific (S,G) entries in the multicast routing table (mroute). With IGMPv1 and v2, a router can only create (\*,G) entries when it receives IGMP JOINS.\
When joining an IGMPv3 group, a host can also specify the sources it wants to receive traffic from (or those that it doesn’t want to receive from). The process is similar to IGMPv2.

#### **Leaving a Group**

Leaving a group in IGMPv3 is also similar to IGMPv2, but IGMPv3 can also enable explicit tracking of each host on the network, thus quickly pruning unneeded traffic from the LAN. To enable explicit tracking, use:

```
R(config-if)# ip igmp explicit-tracking
```

#### **SSM Mapping**

Hosts must support IGMPv3 to be able to participate in SSM. But in some situations you can’t upgrade the clients to support IGMPv3 and you are stuck with them in older versions. Cisco Routers can be used to convert IGMPv1 and IGMPv2 JOINs into SSM/IGMPv3 JOINs. To do this, a router must be configured with a map of available sources for each group that the hosts will join.

```
R(config-if)# ip igmp ssm-map enable
! Enables SSM mapping
R(config-if)#  ip igmp ssm-map static GROUP-ACL SOURCE
! Enables static mapping
R(config-if)#  ip igmp ssm-map query dns
! Enables automatic mapping via DNS.
```

## Router acting as an IGMP Host

A Cisco Router can act as a host by joining a group. Since the traffic is destined to the router, it will be process switched.

```
R(config-if)# ip igmp join-group GROUP
! IGMP v1,v2,v3
R(config-if)# ip igmp join-group GROUP source SOURCE
! only IGMPv3
```

A Cisco Router can also define an interface as a static member of a group, regardless of whether there are receivers or not on that segment. The interface will appear in the OIL of the group and traffic will be fast-switched:

```
R(config-if)# ip igmp static-group {*|GROUP}
! * = all groups
R(config-if)# ip igmp static-group GROUP source {SOURCE|ssm-map}
! Only IGMPv3
```

## Filtering and limiting IGMP Joins

A router can filter incoming IGMP joins and only allow some of them, using:

```
R(config-if)# ip igmp access-group ACL
! For IGMPv1 and v2, IGMP can only be standard
! For IGMPv3, it can also be extended, matching on both group and source
```

You can also limit the number of entries in the mroute table generated by IGMP JOINs. You can do this globally:

```
R(config)# ip igmp limit MAX
```

or on an interface

```
R(config-if)# ip igmp limit MAX [except ACL]
! groups matched by the ACL will not be counted
```

The first limit that is reached is applied.

## Stub multicast routing

The multicast stub router feature is used when you can’t use PIM on a link but you still need to receive multicast traffic for a client that is connected to you. The Stub router will forward IGMP REPORTs to an upstream router which will add this interface into its OIL. To configure the stub router, use:

```
SR(config-if)# ip igmp helper-address IP-ADDR
! configure this on the stub (downstream) router
```

The upstream router should be configured to filter all PIM messages from the downstream Stub Router using one of the following commands:

```
UR(config-if)# ip pim neighbor-filter ACL
! the ACL should deny the Stub router
UR(config-if)# ip pim passive
```


# Inter Domain Multicast

## MSDP (Multicast Source Discovery Protocol)

PIM-DM is not a viable solution to use over Internet, so we will use PIM-SM for interdomain multicast routing. PIM-SM Requires a Randezvous Point. In PIM-SM, when a source starts to transmit, the DR on it’s segment will register the source with the RP.Then, the RP can join the SPT to the source. But what happens when you have the source in one AS and receiver in another one? You could have a single RP for both ASes, but this is a highly improbable since 2 ASes should be under different administration. The better solution is to use MSDP.\
MSDP, defined in [RFC 3618](https://www.ietf.org/rfc/rfc3618.txt), is used between different domains/ASes to signal each other about the sources they have in the domain. In fact, the RP in one AS will signal the RP in the other AS about a new source that wants to transmit. Knowing the source, the RP in the second AS can now join the SPT tree towards the source in the first domain.

MSDP is somewhat similar to BGP because the RP in each domain establishes TCP connections with the RP in the other domains. MSDP uses TCP 639 for the TCP connection.\
When a RP discovers a new source in its domain (via PIM REGISTER), it sends a SOURCE-ACTIVE (SA) message to all MSDP Peers.

When a RP has a (\*,G) entry for the group (meaning a receiver on the shared tree), it will JOIN the SPT tree to the source by sending an (S,G) JOINs.

You can see now that MSDP is used only for exchanging information about existing sources, not for the actual exchange of data.

If there are multiple interconnected domains, a router may receive MSDP information from multiple neighbors. It will perform a special RPF check by checking if the MSDP message that was received came on the same interface as the one that points to the next-hop for reaching the Originator ID, in the BGP table. If also the multicast address-family is used, the information in advertised here is used for the RPF check.

First you must define the MSDP peers:

```
R(config)# ip msdp peer PEER-IP [connect-source INTERFACE] [remomte-as AS-NUMBER]
! configure the MSDP peer and the source interface
```

Additional settings can be configured for each pear:

```
R(config)# ip msdp password peer PEER-IP PASSWORD
! form md5 authentication
R(config)# ip msdp sa-limit PEER-IP LIMIT
! Limit the number of SA messages that it can receive. Prevents DoS
R(config)# ip msdp keepalive PEER-IP KEEPALIVE HOLD-TIME
! Default: KEEPALIVE 60 , HOLD-TIME: 75
R(config)# ip msdp sa-filter {in|out} PEER-IP [list ACL|route-map ROUTE-MAP]
! Restrics the information that is advertised (out) to or is accepted (in) from other peers
! deny entries in the ACL/ROUTE-MAP will be filtered
```

or for the MSDP process:

```
R(config)# ip msdp originator-id INTERFACE
! changing the default Origiantor ID. By default it's the RP address
R(config)# ip msdp redistribute [list ACL]
! Restricts which sources are advertised to other peers
```

### Anycast RP

Anycast RP is an application of multicast inside a domain, that offers load sharing and backup for a RP. Two or more routers are configured with the same /32 Loopback address, and that address is advertised as the RP. The PIM SM routers in the domain will use the Anycast RP that is closest to them. Because the clients and the sources might end up using different RPs, a MSDP session is used to exchange source information between the RPs.\
Fist, define the 2 routers with the same Loopback IP address, and set this address as the PIM RP.\
Then, define the 2 routers as MSPD peers. Make sure you don’t use the same another address as the connect source:

```
R(config)# ip msdp peer PEER-IP [connect-source INTERFACE]
```

This configuration also provides redundancy for the RP.

## MP-BGP for multicast

Multicast only works if the 2 domains are interconnected with links that run PIM. Also, multicast traffic must pass RPF check in order to be forwarded. When performing the RPF check, a router that uses PIM will look into the unicast routing table for the route that points to the destination of the multicast traffic. The unicast routing table may be populated with routes from all routing protocols, including BGP.\
However, what if you need the inter-domain multicast traffic to use another link, different than unicast traffic? Well, you could use static mroutes, but these are local to each router. Anoher option is to use MP-BGP’s **address-family ipv4 multicast**\
What this does, is to exchange unicast routes between routers, but these routes will not be used for unicast routing, instead they will be used for RPF checks when routing multicast traffic.\
Actually, you can have the same routes that are advertised by unicast ipv4 BGP advertised in multicast ipv4 BGP. But since they are used for different functions, there’s no problem. Using this method you can apply different routing policies for multicast and unicast traffic.\
To enable the exchange of routes used for multicast routing, you must activate neighbors in the **address-family ipv4 multicast**. This address-family runs completely independent of the ipv4 unicast address-family, so you will need to configure it appropriately (route-reflectors, advertised networks, redistribution, and so on).

An MP-BGP multicast learned route is preferred over any unicast routes when performing the RPF check.

```
R(config-router)# address-family ipv4 multicast
R(config-router-af)# neighbor NEIGH-ADDR activate
! If needed, originate routes or do other settings in the address-family
R(config-router-af)# network NETWORK mask MASK
```

To check the status of MP-BGP for multicast, use:

```
R# show bgp ipv4 multicast [summary|neighbors NEIGH-ADDR]
```


# IPv6 Multicast

## IPv6 Multicast Addressing FF00::/8

Multicast addresses represent a group of interfaces, just like in IPv4 multicast. The format of a multicast address is:

```
| 8 bits | 4b | 4b |     112 bits   |
+--------+----+----+----------------+
|11111111|0RPT|SCOP|    Group ID    |
+--------+----+----+----------------+
```

Flags:\
The R Flag is used for addresses with an embedded RP.\
The P Flag indicates a multicast address that is assigned based on the network prefix.\
The T Flag is set to 0 if the address is a well-known multicast (assigned by IANA) or 1 if the address is an administratively assigned address

The SCOP field (4 bits, 16 values) contains the scope of the multicast group. Available values are:

* **1**: Interface Local – spans only on a single interface and is used for loopback multicast
* **2**: Link Local – spans on a single link (point-to-point or multi-access)
* **4**: Admin Local – administratively defined
* **5**: Site Local – spans a single site
* **8**: Organizational Local – multiple sites belonging to a single organization
* **E**: Global

IANA keeps the [list of permanently assigned multicast addresses](https://www.iana.org/assignments/ipv6-multicast-addresses/ipv6-multicast-addresses.xml). The most used are:

| Address  | Description       |
| -------- | ----------------- |
| FF02::1  | All Nodes         |
| FF02::2  | All Routers       |
| FF02::5  | All OSPF Routers  |
| FF02::6  | All OSPF DRs      |
| FF02::9  | All RIP Routers   |
| FF02::A  | All EIGRP Routers |
| FF02::D  | All PIM Routers   |
| FF02::12 | All VRRP Routers  |
| FF02::16 | All MLD Routers   |

Other Addresses, like NTP are assigned by IANA but have a variable scope. The address assigned is FF0X::101, where X can be any of the scopes seen earlier.

## IPv6 Layer 2 Addresses

### Ethernet

The frames that carry multicast traffic have the MAC address of the sender as the source and a special address derived from the Layer 3 destination IPv6 as the destination.\
L2 Multicast addresses start with **33:33** followed by the last 32 bits in the IPv6 address.

## MLD

MLD (Multicast Listener Discovery Protocol) replaces IGMP for IPv6 multicast routing. MLDv1 is similar to IGMPv2, while MLDv2 is similar to IGMPv3 (supports Source Specific Multicast).\
To see the interfaces that run MLD, use:

```
R# show ipv6 mld interface
```

## PIM

The default version of PIM used on Cisco routers, PIMv2 also supports IPv6. To enable IPv6 multicast routing, use:

```
R(config)# ipv6 multicast-routing
```

This command also enables PIM on all IPv6 enabled interfaces.\
Only PIM-SM is supported for IPv6 and it works similar to IPv4 PIM.

### PIM Tunnels

Every router creates a tunnel to the RP that it uses to unicast PIM Register messages. These tunnels are automatically created. You can see them using:

```
R# show ipv6 pim tunnels
```

### Rendezvous Point

RP can be configured statically or using BSR. Auto-RP is not supported for IPv6.\
To see a list of configured RPs, use:

```
R# show ipv6 pim range-list
```

#### **Static RP**

To configure a static RP, use:

```
R(config)# ipv6 pim rp-address IPV6-ADDR [bidir]
```

#### **BSR**

To configure a BSR candidate, use:

```
R(config)# ipv6 pim bsr candidate bsr IPV6-ADDR [HASH] [priority PRI]
```

To configure a RP candidate, use:

```
R(config)# ipv6 pim bsr candidate rp IPV6-ADDR [group-list ACL] [priority PRI] [bidir]
```

#### **Embedded RP**

A new way to specify the RP is to embedded it into the multicast group address. The format of the group address is:

```
| 32 bits |      64 bits      | 32 bits  |
+---------+-------------------+----------+
|FF7S:0iLL|     RP Prefix     | Group ID |
S  = SCOPE
i  = Interface ID (4 bits)
LL = prefix Length. Ex: /64 => 0x40
RP Address: RP_Prefix::i/LL
```

E.g. If you want to use a RP with the address 2002:CC1E::9, you should assign the following address to the group: FF7x:0940:2002:CC1E::1. Disecting this address, we get:

* FF7 – marks the embedded RP type
* x – set it to the scope. E.g: for global scope: FF7E:940:2002:CC1E::1. For site local scope: FF75:940:2002:CC1E::1
* 9 – interface ID
* 40 – Hex version of the prefix lenght 0x40 = 64
* 2002:CC1E – prefix of the RP address
* ::1 – ID of the multicast group

Since the RP can be determined on each packet, there is no need for a RP protocol. Still, the RP needs to be configured with its own IP Address as the RP in order to accept registrations:

```
R(config)# ipv6 pim rp-address IPV6-ADDR
```

## Static IPv6 mroutes

IPv6 static multicast routes are added just like normal unicast routes, but with the **multicast** keyword at the end:

```
R(config)# ipv6 route IPV6-SOURCE {IPV6-PREVIOUS-HOP| INCOMING-INTERFACE multicast}
```


# Multicast features on switches

## IGMP Snooping

### Enabling IGMP Snooping

Normally, a switch should forward multicast frames out all ports. It will forward the traffic to all receivers, but it will also forward it to many hosts that don’t need it. To optimize this behavior, Cisco routers implement IGMP Snooping. With IGMP Snooping, a switch listens to IGMP JOINs and LEAVEs to know which ports should receive traffic for a certain multicast destination.\
IGMP Snooping is on by default for all VLANs, If disabled, you can re-enable it with:

```
Sw(config)# ip igmp snooping [vlan VLAN-ID]
```

Snooping works not only with IGMP, but also with CGMP, PIM or DVMRP. If you want to limit learning only for one of these protocols, use:

```
Sw(config)# ip igmp snooping [vlan VLAN-ID] mrouter learn {cgmp | pim-dvmrp}
```

For IGMPv2 you can enable immediate leave support. This is useful when only one host is connected to a port. Therefor there is no need to leave the port in forwarding state if the only host doesn’t need the multicast traffic anymore:

```
R(config)# ip igmp snooping vlan VLAN-ID immediate-leave
```

For IGMPv1 and v2 you can enable Report Suppression. That is when an IGMP Query is sent, only the first Report is allowed, if other reports are being sent by other hosts they will be dropped, because the router will not use them. This feature is on by default, but if disabled you can enable it wit:

```
R(config)# ip igmp snooping report-suppression
```

### Configuring host ports

Usually a switch will detect a port where a PIM router is connected. However, you can configure it statically as a Multicast Router port with:

```
Sw(config)# ip igmp snooping vlan VLAN-ID mrouter interface INTERFACE
```

IGMP hosts are learned by snooping on IGMP messages they send, but you can also add static members to the list of forwarding ports:

```
Sw(config)# ip igmp snooping vlan VLAN-ID static GROUP-ADDR interface INTERFACE
```

To monitor, use:

```
Sw# show ip igmp snooping
```

### IGMP Querier

If there is no router on the VLAN that can send the queries you can enable the IGMP Querier function on the Switch. To do this, you should have a SVI with an IP Address running in that VLAN. If there’s no address on the SVI then a Global IGMP Querier address can be used. To enable IGMP Querier function, use:

```
R(config)# ip igmp snooping querier
R(config)# ip igmp snooping querier address IP-ADDR
! Define a global address to be used by the Querier, if the SVI doesn't have one
```

## Filtering IGMP

### IGMP Profiles

To filter IGMP messages for some multicast groups, first define an IGMP Profile:

```
Sw(config)# ip igmp profile PROFILE-NUMBER
Sw(config-igmp-profile)# permit|deny
!default: deny
Sw(config-igmp-profile)# range GROUP-ADDRESS
```

Then, apply the profile on a L2 interface:

```
Sw(config-if)# ip igmp filter PROFILE-NUMBER
```

### Max Groups

You can also set a maximum number of groups an interface can JOIN, using:

```
Sw(config-if)# ip igmp max-groups MAX
```

When the maximum number of groups is reached, a new IGMP REPORT is dropped by default, but it can also replace an existing GROUP if configured:

```
Sw(config-if)# ip igmp max-groups action {deny | replace}
```

## CGMP – Cisco Group Management Protocol

CGMP is a used between Cisco routers and switches only. With CGMP, a router is able to send information gathered through IGMP to the switch in order to optimize the CAM table for multicast traffic.\
To enable CGMP on a router, use:

```
R(config-if)# ip cgmp
! If used on a L3 switch, apply the config on a L3 interface
```

CGMP messages are sent only by the routers to a well known MAC address (0100.0CDD.DDDD) that the switches listen to and forward out all ports (in order to reach other switches). Through CGMP, a router informs the switches that a specific host (identified by its MAC Address) wants to receive multicast traffic for a specific group (identified by its L2 address 0100.5Exx.xxxx). The switches will then add entries into their CAM tables so that when they receive the multicast traffic (destined for MAC address 0100.5Exx.xxxx) it will only be forwarded out ports where the hosts that joined that group are located (by searching the host MAC address into existing CAM table). This way, there’s no unnecessary traffic sent. A similar mechanism exists to remove entries.

## MVR – Multicast VLAN Registration

MVR feature is somewhat similar to the voice VLAN feature. All multicast traffic for a Group Address is bound to a single VLAN, and receivers send IGMP JOINs to ask for traffic from the multicast VLAN.\
First, we must configure the switch to support MVR:

```
Sw(config)# mvr
! Enables MVR
Sw(config)# mvr group GROUP-ADDR [COUNT]
! the switch will put traffic for this address in the MVR VLAN
! The COUNT will allow MVR to work for several contiguous Group Addresses
Sw(config)# mvr vlan VLAN-ID
! sets the VLAN used for MVR traffic
Sw(config)# mvr mode {dynamic|compatible}
! Dynamic - forwards IGMP Joins to the router
! Compatible - doesn't forward IGMP Joins. The router must be configured to force multicast traffic downstream (default)
```

To configure the source ports:

```
Sw(config-if)# mvr type source
! Source ports must be part of the Multicast VLAN
```

To configure receivers:

```
Sw(config-if)# mvr type receiver
Sw(config-if)# mvr vlan VLAN-ID group [GROUP-ID]
! Statically adds ports to the MVR, without IGMP JOINs
Sw(config-if)# mvr immediate
! Enables Immediate Leave for MVR
```

To monitor, use:

```
Sw# show mvr [members|interface] ...
```


# NAT 101

## Inside, Outside, Local, Global

When defining NAT it is important to understand what Inside/Outside and Local/Global mean. When we use NAT, our router will be at the border between the Inside and the Outside zones. Hosts on the Inside are configured with addresses from the LOCAL Address Space and hosts on the Outside are configured with addresses from the GLOBAL Address Space. The router is the only device that is directly connected to both the Inside and the Outside.

* Hosts
  * **Inside** – hosts on the Inside
  * **Outside** – hosts on the Outside
* Addresses
  * **Local** – addresses used on the Inside (usually private addresses) - Locally Visible
  * **Global** – addresses used on the Outside (usually public addresses) - Globally Visible

## Types of NAT

{% tabs %}
{% tab title="Static NAT" %}
Static NAT is the NAT operation where we have a 1:1 relationship between the Local IP and the Global IP. This is typically used when you want to allow outside hosts to initiate connections to inside hosts

The Static NAT makes a permanent entry in the NAT table that links the local address to the global address
{% endtab %}

{% tab title="Dynamic NAT" %}
**Dynamic NAT** is the NAT operation where we have many to many relationships between local IP addresses and global IP Addresses.&#x20;

Dynamic NAT entries may timeout from the NAT table so the same global address can later be associated with a different local address without any change of the config.

When you think about it, Dyamic NAT means that the local IP addresses will be associated with one global IP address at the time but once all global IP addresses available are consumed, a new host that wants to communcate with the outside will not be able to get a global IP assigned.

One form of Dynamic NAT is Dyanmic NAT with **Oveload** or **PAT** (Port Address Translation). In this case multiple local addresses can share one global IP Address (or sometimes a smaller pool of addresses). This is achieved by also translating the L4 information so that a unique association is formed between the local source and port and the global source and port.
{% endtab %}
{% endtabs %}

## Classic NAT

### 1. Choose the sides: Inside and Outside

The basic idea with NAT is to change the Layer 3 address (either source or destination) of a packet. (You can change also the Layer 4 information, but this will come up later).

In most common scenarios you have 2 sides, called Inside and Outside and you will have to chose were your interfaces belong to.

```
R(config)# interface INSIDE-INTERFACE
R(config-if)# ip nat inside
R(config)# interface OUTSIDE-INTERFACE
R(config-if)# ip nat outside
```

In simple scnearios, you will have at least one interface on each side, but you can have multiple inside and outside interfaces. It all comes down to how the interface that is receiving the packet is configured for NAT which drives the order of operations:

1. When packets arrive on an **Inside** interface&#x20;
   1. they are **first routed** (based on destination IP address)&#x20;
   2. **then translation** occurs for either the source or the destination address.&#x20;
2. When packets arrive on the **Outside** interface,&#x20;
   1. they are f**irst translated** (either the source or the destination address)&#x20;
   2. **then** they are **routed** based on the destination address (which may have just been translated).

### 2. Create NAT bindings

This action is different depending on the type of NAT you will use but they are all configured from the global config mode:

```
R(config)# ip nat {inside|outside} ...
```

You will find more details about the options further down

### 3. Monitor NAT bindings

To monitor the bindings that the router knows about use the below command. It will show the Inside global, Inside local, Outside local and Outside global values in it's NAT table

```
R# show ip nat translations
R# show ip nat statistics
```

NAT translations that are not in use can be cleared from the NAT table with&#x20;

```
R# clear ip nat translation *
```

Of course, this doesn't apply to static entries which will appear in the NAT table even when there is no traffic.

## NAT Inside Source

When translating the Inside Source, we tell the router to translate the Source address of the packets that arrive on the Inside. So we first need to define an inside interface:

```
R(config)# interface INSIDE-INTERFACE
R(config-if)# ip nat inside
```

When the packet arrives, the router will first route it (based on the destination – OUTSIDE\_GLOBAL) and then it will translate the source IP. An entry is created into the NAT Translation table that maps the INSIDE\_LOCAL address to an INSIDE\_GLOBAL address.\
When the Outside Host replies, it sends return packets to the INSIDE\_GLOBAL destination address. If you want to get this traffic to the Inside Host, you have to make sure you receive it on an Outside interface

```
R(config)# interface OUTSIDE-INTERFACE
R(config-if)# ip nat outside
```

When the outside interface receives a packet, it will first translate it (based on the same entry into the NAT Translation table) and then it will route it (based on the new destination – INSIDE\_LOCAL

![NAT inside source](/files/SZXkcMm4ZIKLM8NPcSnj)

There are different options when implementing inside source NAT:

### Host Static NAT

Used to map one INSIDE\_LOCAL address to one INSIDE\_GLOBAL address:

```
R(config)# ip nat inside source static INSIDE_LOCAL INSIDE_GLOBAL
```

This command maps the INSIDE\_LOCAL address to the INSIDE\_GLOBAL address when performing source inside NAT. All IP packets sourced by the INSIDE\_LOCAL address get their source translated to the INSIDE\_GLOBAL, and all IP packets destined to the INSIDE\_GLOBAL address get their destination translated to INSIDE\_LOCAL.\
You can’t map the same INSIDE\_GLOBAL address to multiple INSIDE\_LOCAL addresses, or the other way around. Since this is a 1-to-1 mapping that is always present in the NAT table, connection can be initiated by any side (Inside and Outside).

### Port Static NAT

This configuration is similar with the one before, except that instead of translating all IP traffic, only traffic for 1 UDP or TCP port is translated. This allows the router to use the same INSIDE\_GLOBAL address for multiple INSIDE\_LOCAL addresses, as long as they use different GLOBAL\_PORTs

```
R(config)# ip nat inside source static {tcp|udp} INSIDE_LOCAL LOCAL_PORT INSIDE_GLOBAL GLOBAL_PORT
```

### Network Static NAT

Usually used when dealing with overlapping networks, this command allows translating the INSIDE\_LOCAL network to an INSIDE\_GLOBAL network. However, both networks have to have the same NETMASK.

```
R(config)# ip nat inside source static network INSIDE_LOCAL_NETWORK INSIDE_GLOBAL_NETWORK [/PREFIXLEN|NETMASK]
```

### Dynamic NAT

A dynamic map is used to enable translation for multiple hosts on the inside, and uses a many-to-many mapping, where multiple INSIDE\_LOCAL addresses are mapped to multiple INSIDE\_GLOBAL addresses.

For this, we must first define a pool of addresses to use as INSIDE\_GLOBAL and an ACL for the INSIDE\_LOCAL addresses:

```
R(config)# ip nat pool INSIDE_GLOBAL_POOL START-IP END-IP {netmask NETMASK|prefix-length LEN}
R(config)# access-list INSIDE_LOCAL_ACL permit INSIDE_LOCAL_ADDR [WILDCARD]
! It also works with an extended ACL
```

Then, enable Dynamic NAT with:

```
R(config)# ip nat inside source list INSIDE_LOCAL_ACL pool INSIDE_GLOBAL_POOL
```

Dynamic translations generate entries in the NAT table only when traffic comes on the Inside interface and is routed to the Outside Interface. This is when the router takes one address from the INSIDE\_GLOBAL\_POOL and maps it to the INSIDE\_LOCAL address that sourced the traffic. Traffic on the reverse path is dropped if there is no translation generated by traffic on the inside. If the translation times out, the INSIDE\_GLOBAL address is returned to the pool. But with enough active translations you can empty the INSIDE\_GLOBAL\_POOL, and conversations from new hosts on the INSIDE will not be allowed.

### Dynamic NAT with overload – PAT

When using Dynamic NAT without overload you could only have a small number of translations limited by the size of the INSIDE\_GLOBAL\_POOL. If the pool was empty, conversations from new hosts on the INSIDE were not allowed. With overloading, the source port is dynamically translated into a unique global port, so the router differentiate between inside hosts and uses fewer (even one) INSIDE\_GLOBAL addresses for many hosts on the inside.\
For this, we must again define a pool of addresses to use as INSIDE\_GLOBAL and an ACL for the INSIDE\_LOCAL addresses:

```
R(config)# ip nat pool INSIDE_GLOBAL_POOL INSIDE_GLOBAL_START INSIDE_GLOBAL_END {netmask NETMASK|prefix-lenght LEN}
R(config)# access-list INSIDE_LOCAL_ACL permit INSIDE_LOCAL_ADDR [WILDCARD]
! It also works with an extended ACL
```

Then, enable NAT with overloading:

```
R(config)# ip nat inside source list INSIDE_LOCAL_ACL pool INSIDE_GLOBAL_POOL overload
```

### NAT Outside Source

This type of NAT is used when you want the Outside Hosts to appear as Local Hosts. However, there are some problems to take care of so this setup is not that widely used.

![NAT outside source](/files/s2QY9pDesDM5AKz3QgPd)

With Outside Source NAT, when a packet sourced by OUTSIDE\_GLOBAL and destined to an INSIDE\_LOCAL address arrives on the Outside interface, it is first translated and then routed. This means that first we translate the source address (OUTSIDE\_GLOBAL to OUSIDE\_LOCAL) and then we should route based on the destination IP. The destination is the INSIDE\_LOCAL address so traffic should arrive to the Inside Host. However, return traffic will be packets sourced by the INSIDE\_LOCAL address and destined to the OUTSIDE\_LOCAL address. This is where things get tricky, because the destination is from the Local Address Space and we should send the packet back on the Inside interface. And this is exactly what happens unless we define a static route that forces the router to forward the packet on the Outside interface. You can see this by looking at the command **debug ip nat** and you will see that the traffic is returned to the incoming interface. One way to resolve this is by adding a more specific static route that points to the actual destination.\
Another way is to perform another Static NAT Translation for the return traffic. You will need a translation that takes the destination address\
You can make the router add the static route by itself if you use the **add-route** keyword but personally I have seen it fail, so a better way to do it is by using static routes.\
In some funky configuration you can even make another NAT Outside Source translation on the return path, which will result in have 2 outside interface defined and no inside interface.\
Now, based on the requirements, you may have to use both inside source and outside source translation, for example when dealing with [overlapping networks](https://nyquist.eu/nat-for-overlapping-networks/) and you may even have to use more static routes.\
As with inside NAT we have multiple ways of configuring the outside NAT

### Host Static NAT

```
R(config)# ip nat outside source static OUTSIDE_GLOBAL OUTSIDE_LOCAL
```

### Port Static NAT

```
R(config)# ip nat outside source static {tcp|udp} OUTSIDE_GLOBAL GLOBAL_PORT OUTSIDE_LOCAL LOCAL_PORT
```

### Network Static NAT

```
R(config)# ip nat outside source static network OUTSIDE_GLOBAL_NETWORK OUTSIDE_LOCAL_NETWORK {/PREFIXLEN|NETMASK}
```

### Dynamic NAT

```
R(config)# ip nat pool OUTSIDE_LOCAL_POOL START_IP END_IP {netmask NETMASK | prefix-length LEN}
```

```
R(config)# ip nat outside source list OUTSIDE_GLOBAL_ACL pool OUTSIDE_LOCAL_POOL 
```

## Inside Destination Address Translation – NAT for Server Load Balancing

Use it to redirect traffic for a SERVER-IP to a pool of real servers behind NAT. The process is somewhat similar to the Inside Source Address Translation, only this time traffic from Outside generates entries in the NAT table so the perspective changes:\
First define a pool that contains the INSIDE\_LOCAL servers, and an ACL that matches the OUTSIDE\_GLOBAL addresses of the server:

```
R(config)#ip nat pool REAL_SERVER_POOL START_IP END_IP {netmask NETMASK | prefix-length LEN} type rotary
R(config)# ip access-list GLOBAL_SERVER_ACL GLOBAL_ADDR [WILDCARD]
```

Then, enable the translation with:

```
R(config)# ip nat inside destination list GLOBAL_SERVER_ACL pool REAL_SERVER_POOL
```

The router will choose a new address from the REAL\_SERVER\_POOL, whenever a new Outside Host tries to access an IP from the GLOBAL\_SERVER\_ACL

## NAT Virtual Interface (NVI)

This feature removes the need to  define inside and outside interfaces. Instead you enable the interfaces for NVI and then the NAT order of operations changes such that routing operations happen twice:

1. Routing
2. NAT operation
3. Routing again with the new translated addresses

```
R(config)# access-list 1 permit 10.0.0.0 0.0.0.255
R(config)# ip nat source list 1 interface e0/1 overload
R(config)# int e0/0
R(config-if)# ip nat enable
R(config-if)#int e0/1
R(config-if)# ip nat enable 
```

What NVI helps with is allowing NATted trafic to flow between two inside interfaces.


# NAT for Overlapping Networks

When we have 2 networks with overlapping addresses, chances are it’s not going to work. Unless, you use NAT.\
The situation we have to deal is can be seen in the next diagram

![NAT for Overlapping Networks – Example Topology](/files/W2PCoxWL42ybE7mZQScU)

We have 2 routers, Router 1 and Router 2, connected via the 12.12.12.0/24 subnet. Each router has a LAN interface on the 10.0.0.0/24 subnet (the overlapping networks).\
How can we communicate between Server 3 and Server 4 without changing the addresses on each router? Let’s take it from the top.

The starting configurations are:

```
!Router 1:
R1(config)# int Fa0/1
R1(config-if)# ip address 10.0.0.1 255.255.255.0
R1(config-if)# int Fa0/0
R1(config-if)# ip address 12.12.12.1 255.255.255.0
!Router 2:
R2(config)# int Fa0/1
R2(config-if)# ip address 10.0.0.2 255.255.255.0
R2(config-if)# int Fa0/0
R2(config-if)# ip address 12.12.12.2 255.255.255.0
```

## Option 1 – NAT on both routers

Translate Server3’s IP address into 13.13.13.3 and Server4’s IP address int 24.24.24.4. Server 3 will use 24.24.24.4 to access Server 4, and Server 4 will use 13.13.13.3 to access Server 3. Each router will need to be aware of the inside global addresses used on the other router.

```
! On R1
R1(config-if)# int Fa0/1
R1(config-if)# ip nat inside
R1(config-if)# int Fa0/0
R1(config-if)# ip nat outside
R1(config-if)# exit
! Translate inside local 10.0.0.3 to inside global 13.13.13.3
R1(config)# ip nat inside source static 10.0.0.3 13.13.13.3
! Add a route towards the inside global subnet used on the other router:
R1(config)# ip route 24.24.24.0 255.255.255.0 12.12.12.2
!On R2
R2(config-if)# int Fa0/1
R2(config-if)# ip nat inside
R2(config-if)# int Fa0/0
R2(config-if)# ip nat outside
R2(config-if)# exit
! Translate inside local 10.0.0.4 to inside global 24.24.24.4
R1(config)# ip nat inside source static 10.0.0.4 24.24.24.4
! Add a route towards the inside global subnet used on the other router:
R1(config)# ip route 13.13.13..0 255.255.255.0 12.12.12.1
```

This was simple, but what if we have more than one IP addresses that must talk to each other? Adding static NAT entries for each host would not be much fun. The first option that comes to mind is to use dynamic NAT instead of the static NAT. Unfortunately, translations for inside source dynamic NAT are only created when the inside host initiates the traffic. We would have a situation, where most of the time the connection would not work, but some times it would. Can you guess when?\
The answer is when both hosts initiate a connection at about the same time (inside the timeout interval) and translation rules are created on both routers. Otherwise, the traffic would reach the router on the other end, but it would not know how to send it to the inside host.

Depending on the type of applications you can use Static NAT on one side and dynamic NAT on the other, but not both dynamic.

## Option 2 – NAT on one side only

Sometimes you do not have access to all devices in the network. Maybe R2 belongs to another company. We have to do it all in our R1 router. This can be done if we change both the source and the destination address in the packet:

```
R2(config)# int Fa0/0
R1(config-if)# ip nat outside
R1(config-if)# int Fa0/1
R1(config-if)# ip nat inside
R1(config-if)# exit
R1(config-if)# ip nat inside source static 10.0.0.3 13.13.13
R1(config-if)# ip nat outside source static 10.0.0.4 24.24.24.4
```

Again, depending on the side that initiates the connection you can use one static NAT and one dynamic NAT:

### Server 3 initiates

```
R(config)# ip nat inside source list 3 pool POOL3
R1(config)# ip nat outside source static 10.0.0.4 24.24.24.4
R1(config)# ip nat pool POOL3 13.13.13.1 13.13.254 prefix-length 24
R1(config)# access-list 3 permit 10.0.0.0 0.0.0.255
```

Server 3 will access Server 4 using 24.24.24.4 address

### Server 4 initiates

```
R1(config)# ip nat inside source static 10.0.0.3 13.13.13.3
R1(config)# ip nat outside source list 4 pool POOL4
R1(config)# ip nat pool POOL4 24.24.24.24.1 24.24.24.24.254 prefix-length 24
R1(config)# access-list 4 permit 10.0.0.0 0.0.0.255
```

Server 4 will access Server 3 using 13.13.13.3 address


# ACLs 101

An ACL contains one or more ACEs (Entries) that permit or deny traffic and have an implicit deny any at the end.

## Numbered ACLs

### Standard ACLs

```
R(config)# access-list ACL-NUMBER {permit|deny} {IP-ADDRESS [WILDCARD] | any} [log]
! ACL-NUMBER: 1-99, 1300-1999
! when the wildcard is missing, a default of 0.0.0.0 is considered
! any <=> IP-ADDRESS 255.255.255.255
! log = adds an entry in the log (one entry every 5 minutes)
R(config)# access-list ACL-NUMBER remark COMMENT
```

You cannot edit one individual entry in a numbered ACL. The ACL must be deleted and re-created.

### Extended ACLs

```
R(config)# access-list ACL-NUMBER {permit|deny} PROTOCOL {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} [OPTIONS] [log|log-input]
! ACL-NUMBER: 100-199, 2000-2699
! PROTOCOL = ip, tcp, upd, protocol number, etc
! any <=> IP 255.255.255.255
! host IP <=> IP 0.0.0.0
! log = adds an entry in the log (one entry every 5 minutes)
! log-input = adds additional info to the log (input interface, source MAC)
! OPTIONS: dscp, precedence, tos, IP Options, fragments, ttl...
R(config)# access-list ACL-NUMBER remark COMMENT
```

#### **Established**

One option for TCP traffic is to allow only packets that are part of an established connection. Packets that are part of an established TCP connection have the ACK or RST bit set. When using the **established** keyword at the end of TCP extended ACL, you can match only these packets:

```
R(config-ext-acl)# {permit|deny} tcp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} established
```

This is useful when you don’t want one side to initiate the connection, but you need it to be able to respond to connections initiated from the other side.

#### **Matching Tips**

Here are some tips for matching traffic with extended ACLs:

```
! Match RIP:
R(config-ext-acl)# {permit|deny} udp {any|SRC-IP SRC-WILDCARD} any eq {520|rip}
! Match EIGRP
R(config-ext-acl)# {permit|deny} eigrp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD}
! Match OSPF
R(config-ext-acl)# {permit|deny} ospf {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD}
! Match BGP
R(config-ext-acl)# {permit|deny} tcp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} eq {179|bgp}
! Match LDP
R(config-ext-acl)# {permit|deny} tcp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} eq 646
R(config-ext-acl)# {permit|deny} udp {any|SRC-IP SRC-WILDCARD} any eq 711
! Match TDP
! Match LDP
R(config-ext-acl)# {permit|deny} tcp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} eq 711
R(config-ext-acl)# {permit|deny} udp {any|SRC-IP SRC-WILDCARD} any eq 646
! Match FTP
R(config-ext-acl)# {permit|deny} tcp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} range 20 21
!Match ping
R(config-ext-acl)# {permit|deny} icmp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} echo
R(config-ext-acl)# {permit|deny} icmp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} echo-reply
!Match traceroute
R(config-ext-acl)# {permit|deny} icmp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} time-exceeded
R(config-ext-acl)# {permit|deny} icmp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} port-unreachable
R(config-ext-acl)# {permit|deny} udp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} range 33434 33464
! Match Path MTU Discovery
R(config-ext-acl)# {permit|deny} icmp {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} packet-too-big
```

To see a list of the most used well-known ports, use:

```
R# show ip port-map
```

## Named ACLs

### Standard ACLs

```
R(config)# ip access-list standard ACL-NAME
R(config-std-nacl)# {permit|deny} {IP-ADDRESS [WILDCARD] | any} [log]
R(config-std-nacl)# remark COMMENT
```

### Extended ACLs

```
R(config)# ip access-list extended ACL-NAME
R(config-ext-nacl)# [seq-number] {permit|deny} PROTOCOL {any|SRC-IP SRC-WILDCARD} {any|DST-IP DST-WILDCARD} [OPTIONS] [log|log-input]
!seq-number can be used to edit an ACL entry or to insert one entry into an existing ACL
!if seq-number is missing, the ACE will be appended with a seq-number value of max-seq-number+10
R(config-ext-nacl)# remark COMMENT
```

Since ACEs can have a custom seq-numbere, they can be re-sequenced to allow other insertions:

```
R(config)# ip access-list resequence ACL-NAME SEQ-START INCREMENT
```

## Using ACLs

### Filter traffic on an interface

```
R(config-if)# ip access-group ACL {in|out}
```

### Limit CLI access

```
R(config-line)# access-class ACL {in|out}
! ACL - can only be standard ACL
! in - ACL applies to inboud connections
! out - ACL applies to outbound connections. ACL matches destination address
```

## Fragments

By default, an ACL without the fragments keyword works in the following manner:

* If the ACL contains only Layer 3 information (SRC-IP, DEST-IP) the entry is applied to nonfragmented packets, initial fragments and non-initial fragments
* If the ACL contains both Layer 3 and Layer 4 information (SRC-IP, DEST-IP, Layer 4 protocol):
  * The entry is applied to nonfragmented packets and initial fragments
  * If the entry is a permit, it will be applied to non-initial fragments, but since the Layer 4 information is missing, it will only match on Layer 3 information, and will ignore Layer 4 information.
  * If the entry is a deny, it will be ignored for non-initial fragments, and the next ACE is processed

When using the **fragments** keyword, the entry is applied only to non-initial fragments. Non-fragmented packets and initial fragments are not matched. This keyword cannot be used on entries that use Layer 4 information.

## Logging ACLs

You must use the log keyword at the end of an ACE to enable generation of syslog messages when the ACE is hit. One message will be generated every 5 minutes.\
You can set the number of hits that will generate a log entry using:

```
R(config)# ip access-list log-update threshold HITS
```

For extended ACLs, you can use log-input keyword to add more information to the log (input interface, source MAC). You can add custom tags to the logs by using the log keyword followed by the tag, or you can generate automatic hashtags for each ACE by setting:

```
R(config)# ip access-list logging hash-generation
```

These hashtags will identify the ACE that generated them in the log output.


# ACLs 102

## Time-based ACLs

Define the time range:

```
R(config)# time-rage TIME-RANGE
R(config-time-range)# periodic DAYS-OF-WEEK HH:MM to [DAYS-OF-WEEK] HH:MM
! adds a recurring time to the time-range
! DAYS-OF-WEEK: daily (M-S), weekdays(M-F), weekend(S,S), Monday, Tuesday, ...
R(config-time-range)# absolute [start TIME DATE][end TIME DATE]
! adds an absoulte time to the time-range
! Only one absolute time is permitted in one time-range
```

Add the time-range to the ACL:

```
R(config)# access-list ACL {permit|deny} ... time-range TIME-RANGE
R(config-std-nacl)# {permit|deny} ... time-range TIME-RANGE
R(config-ext-nacl)# {permit|deny} ... time-range TIME-RANGE
```

## Reflexive ACLs

A reflexive ACL is used to permit outgoing traffic that was originated on one side of the connection (inside) and allow the returning packets from the other side (outside), but to deny traffic that was originated from the other side (outside). You can only use an extended named ACL to implement Reflexive ACLs

```
!Define the ACL that will permit traffic on the Inside:
R(config)# ip access-list extended ACL-OUTGOING
R(config-ext-nacl)# permit PROTOCOL SRC-IP WILDCARD DST-IP WILDCARD reflect REFLECT-NAME [timeout seconds]
! Defines the ACL that will permit traffic on the Outside:
R(config)# ip access-list extended ACL-INCOMING
R(config-ext-nacl)# evaluate REFLECT-NAME
```

The default timeout for dynamic entries in a reflexive ACL si 5 minutes but this can be changed per ACE or globally:

```
R(config)# ip reflexive-list timeout SEC
```

Usually, the ACL that matches outgoing traffic is set on the inside interface, while the ACL that evaluates the reflexive entries is set either on the outside interface on the incoming direction, or on the inside interface on the outgoing direction

```
! Apply OUTGOING-ACL
R(config)# interface INSIDE
R(config-if)# ip access-group ACL-OUTGOING in
! Apply INCOMING-ACL
R(config)# interface OUTSIDE
R(config-if)# ip access-group ACL-INCOMING in
! OR
R(config)# interface INSIDE
R(config-if)# ip access-group ACL-INCOMING out
```

## Dynamic ACLs – Lock-and-Key

This feature allows an IOS router do dynamically add ACEs in an ACL, in order to allow traffic for specific users. Users that need to pass traffic that is normally blocked by an ACL, can use telnet to logon to the router which then will dynamically add entries to the ACL in order to let them pass the filter.\
First, configure the ACL and include a dyanmic template:

```
R(config-ext-acl)# {permit|deny}...
R(config-ext-acl)# dynamic DYNAMIC-NAME [timeout MIN] {permit|deny} ...
```

The ACL should be applied on an interface but the dynamic ACEs will not be used in the ACL until a users authenticates itself using telnet. We will set this using the autocommand setting, defined on a line or on a username.

```
! Define user autocommand
R(config)# username USER autocommand access-enable [host] [timeout MIN]
! Define line autocommand
R(config-line)# autocommand access-enable [host] [timeout MIN]
```

The access-enable command will add the dynamic entries in the ACL. When using the **host** keyword, the ACE will only allow traffic from the host that connected via telnet. Otherwise, it will allow traffic from anybody.

Usually, you enable dynamic extension of the timeout period. If you login again, the timeout period is extended for each new login, otherwise the dynamic entries will be deleted from the ACE after they expire. To enable this feature, use:

```
R(config)# access-list dynamic-extend
```

You can manually clear dynamic entries using:

```
R# clear access-template ACL [DYNAMIC-NAME]
```


# Cisco IOS Firewall

## CBAC – Context Based Access Control

CBAC allows examination of traffic at the Application Layer, not just Layer 3 or Layer 4 as in ACLs. It can maintain session information and create temporary openings to allow return traffic for permissible sessions.\
CBAC maintains a state table both for TCP and UDP (aproximated state – since the service is connectionless). Packets in the return traffic must match the information in the state table to be allowed.\
CBAC is set on an interface in one direction (incoming or outgoing) and will only inspect those packets that passed the ACL in that direction. (see [Order of operations](https://nyquist.eu/routing-order-of-operations/)).\
when traffic originates in the direction that the inspection rule is applied, it will generate shortcuts that will bypass any ACLs in the opposite direction, regardless of the interface. This means an inspect on the incoming traffic of the inside interface will allow returning packets on the outside (even if there is an ACL that denies them). Also, if the inspection is configured on the outside interface, in the outgoing direction, it will still allow return traffic.

To define the CBAC inspection rule, use:

```
R(config)# ip inspect name INSPECTION-NAME PROTOCOL [alert {on|off} [audit-trail {on|off}] [timeout SEC]
! alert: generates syslog messages
! audit-trail: generates more verbose messages
```

Apply the inspection rule on an interface:

```
R(config-if)# ip inspect INSPECTION-NAME
```

## TCP Intercept

TCP intercept is used to prevent servers from TCP SYN-flood attacks. When this type of attacks occur, an attacker sends multiple TCP SYN packets to a server, which should try to respond with a SYN-ACK and then keep this state information until the sender responds with an ACK and the three-way-handshake is completed. When the number of SYN packets received is high enough, the server may start dropping legitimate connections. The router can be configured to prevent this using the TCP Intercept feature.\
There are two modes of operation: active and passive. In the default active mode (intercept), the router will intercept the TCP SYN packets, and respond with a SYN-ACK on behalf of the server. Only when it receives an ACK back, it will create a three-way-handshake with the server and connect the 2 sessions with each other, thus stopping excess SYNs from reaching the server.\
In the passive mode(watch), the router forwards the SYN to the server but waits a limited amount of time (default:30 sec) for the three-way-handshake to complete, before it sends a RESET to the server to clear the connection.

```
! Traffic passing the following ACL will be intercepted.
! Usually matches the destination server
! Could be used not to intercept some known sources
R(config)# ip tcp intercept list ACL
! Define intercept mode: active (intercept) or passive (watch)
R(config)# ip tcp intercept mode {intercept | watch}
! Define timers
R(config)# ip tcp intercept {watch-timeout| finrst-timeout | connection-timeout} SEC
```

The TCP Intercept also has an aggressive mode in which any new connection will generate the drop of an old connection. This aggressive mode is automatically enabled based on a couple of thresholds regarding total number of incomplete connections or number of connections in the last minute.

```
R(config)# ip tcp intercept max-incomplete low LOW-VAL high HIGH-VAL
R(config)# ip tcp intercept one-minute low LOW-VAL high HIGH-VAL
R(config)# ip tcp interce drop-mode {oldest | random} 
```

## Unicast RPF (Reverse Path Forwarding)

Unicast RPF uses the same mechanism as the [Multicast RPF used by PIM](https://nyquist.eu/multicast-101/#2_Reverse_Path_Forwarding) to verify that a packet arrived on the interface with the best route pointing to the source. If it did, then the packet is processed, if it didn’t, then it is dropped.

```
! CEF must be enabled
R(config)# ip cef
R(config)# interface INTERFACE
R(config-if)# ip verify unicast reverse-path [list ACL]
! spoofed packets that are permited by ACL are permited but can be logged
! spoofed packets that are denied by the ACL are dropped but can be logged
```

For multihomed environments, a loose version of uRPF can be enabled, where the packet does not need to enter a specific interface. A route to the source must exist in the routing table, otherwise the packet will be dropped. To enable the loose-mode uRPF, use:

```
R(config-if)# ip verify unicast source reachable-via any
```


# Zone Based Firewall

## Basics

A Zone Based Firewall uses the same inspection engine as CBAC, but works with security zones, not with individual interfaces.\
A zone groups multiple interfaces together. By default, traffic is allowed between interfaces in the same zone, but is not allowed between interfaces in different zones. You can define zone pairs which are zones that can send traffic to each other.\
A special zone is automatically created for traffic destined to or generated by the router itself. This special zone called “self” allows traffic to/from it to the other interfaces.\
Zone Based Policy Firewall uses a syntax similar to [MQC used in QoS](https://nyquist.eu/qos-101-classifying-and-marking/#2_MQC_Modular_QoS_CLI)

## Configuration

### Define the Policy

```
! Create Class 
R(config)# class-map type inspect [mathc-any | match-all] CLASS
R(config-cmap)# match {access-group ACL | protocol PROTOCOL | class-map CHILD-CLASS}
! Create Policy
R(config)# policy-map type inspect POLICY
R(config-pmap)# class type inspect CLASS
! Define action:
R(config-pmap-c)# inspect PARAMETER-MAP
! Enables the CBAC engine
R(config-pmap-c)# police rate CIR burst BC
! Optional, sets policing
R(config-pmap-c)# drop [log]
! Drops the packets
R(config-pmap-c)# pass
! Allows packets to be forwarded
R(config-pmap-c)# urlfilter URL-PARAM-MAP
```

A parameter map is used to define the parameters of the inspection engine or of the urlfilter engine

```
! Define the parameter-map for inspection:
R(config)# parameter-map type inspect [PROTOCOL] {PARMETER-MAP | global | default}
! PROTOCOL specific options can be set inside the parameter-map
! > Start defining the Inspection parameters:
R(config-profile)# ?
! Define the parameter-map for URL Filtering:
R(config)# paramter-map type urlfilter URL-PARAM-NAME
! > Start defining the URL filtering params
R(config-profile)# ?
```

### Create the zones

```
R(config)# zone security ZONE-NAME
R(config-sec-zone)# description DESCRIPTION
```

### Add interface to a zone

```
R(config-if)# zone-member security ZONE-NAME
```

### Create Zone Pairs and attach policy

```
R(config)# zone-pair security ZONE-PAIR-NAME source {ZONE|self} destination {ZONE|self}
! Attach policy to zone
R(config-sec-zone-pair)# service-policy type inspect POLICY
```

Be aware that when you add an interface to a zone, by default, it will drop all traffic that is destined to other zones, unless there is a policy that permits it.\
One simple mistake is that the policy is only applied on one zone-pair (from interface A to interface B), but no policy is applied on the return direction (from interface B to interface A)\
This may not be harmful when you use Inspection, because inspected traffic will automatically be allowed on the return path, but if you use just **pass**, return traffic will be dropped if no service-policy that explicitly allows it is placed on the return path.


# AAA 101

## Enabling AAA new-model

AAA stands for Authentication, Authorization and Accounting. Authentication is the process of identifying users based on some credentials (passwords, digital certificates, tokens). Authorization is the process of allowing an authenticated user to access specific services or a specific level of administration, while accounting is the process of tracking and logging a user’s actions while authenticated.\
To enable the new-model of AAA implementation on a Cisco device, use:

```
R(config)# aaa new-model
```

## Methods

Authentication methods can be grouped in 2 categories: Group methods and Non-Group methods. Group methods include protocols such as RADIUS, TACACS+ or Kerberos. These protocols require an external server that handles authentication requests. The non-group methods include local login (usernames defined locally), enable passwords or line passwords.

### RADIUS

RADIUS is now an industry standard and runs on UDP 1812 and UDP 1813. A RADIUS packet encrypts only the password field.

To define a radius server that will process AAA requests, use the following commands:

```
R(config)# radius-server host {HOST|IP-ADDR} [auth-port PORT] [acct-port PORT] [timeout SEC] [retransmit RETRIES] [key KEY] [alias {HOST|IP-ADDR}]
! Configures a radius server
```

All Radius servers are by default part of the **radius** group, but custom groups that include subsets of the radius group can be configured:

```
R(config)# aaa group server radius RADIUS-GROUP
R(config-sg-radius)# server IP-ADDR [auth-port PORT] [acct-port PORT]
```

### TACACS+

TACACS+ is a Cisco proprietary protocol that runs on TCP 49. With TACACS the entire packet body is encrypted. TACACS+ is configured in a similar way as RADIUS:

```
R(config)# tacacs-server host {HOST|IP-ADDR} [single-connection] [port PORT] [timeout SEC] [key KEY] 
```

All TACACS+ servers are by default part of the **tacacs+** group, but custom groups that include subsets of the tacacs+ group can be configured:

```
R(config)# aaa group server tacacs+ TACACS+GROUP
R(config-sg-tacacs)# server {HOST|IP-ADDR}
```

### Local and Local-case

Just define the usernames on the local router:

```
R(config)# username USER [privilege {1-15}] [password PASS | secret SECRET]
```

The **local** method use the case-insensitive users locally defined, while the **local-case** method uses case-sensitive values.

### Enable Password

```
R(config)# enable [password PASS | secret SECRET]
```

### Line Passwords

Line password uses the passwords defined on each line:

```
R(config-line)# password PASS
```

## Authentication

### Authentication lists

An authentication list consists of one or more authentication methods that are used, in order, for authentication. The first method in the list is used and if the result is an accept or a reject, the other methods are not used. They are used only if the previous methods are not available (like a server response timeout).\
Authentication lists can be defined for several features that may require authentication. A default list is already applied for each feature, but other lists can be defined, or the default list can be populated.

```
R(config)# aaa authentication LIST-TYPE {default|LIST-NAME} METHODS-LIST
! Some LIST-TYPEs can be:
  dot1x            Set authentication lists for IEEE 802.1x.
  enable           Set authentication list for enable (privileged mode)
  login            Set authentication lists for logins.
  ppp              Set authentication lists for ppp.
! METHODS LIST can include:
  enable       Use enable password for authentication.
  group {GROUP-NAME|radius|tacacs+}        Use Radius/Tacacas+ Server-group
  krb5         Use Kerberos 5 authentication.
  krb5-telnet  Allow logins only if already authenticated via Kerberos V Telnet.
  line         Use line password for authentication.
  local        Use local username authentication.
  local-case   Use case-sensitive local username authentication.
  none         NO authentication
```

Unless a LIST-NAME is specified, the router uses the *default* lists, but the default configuration of the default lists doesn’t show up in the configuration (by default). If you define the default lists yourself, they will show up in the config.\
If you don’t, the router uses the *Permanenet* lists, which are predefined.

```
! For PPP and VTY line authenticaion, it uses the Permanent Local list, which is similar to:
R(config)# aaa authentication ppp default local
! For Console line authentication, it uses the Permanent None list, which is similar to:
R(config)# aaa authentication ppp default none
! For Enable authentication, it uses the Permanent Enable list, which is similar to:
R(config)# aaa authentication enable default enable
```

### Applying authentication lists

#### **Line Authentication (console, vty)**

```
R(config)# line {con | vty LINE-START [LINE-END]}
R(config-line)# login authentication {default|LOGIN-LIST-NAME}
! attach the LOGIN-LIST-NAME to a line
```

#### **PPP Authentication**

```
R(config-if)# ppp authentication {chap|pap|...} {default | PPP-LIST-NAME}
! authenticate using the PPP-LIST-NAME
```

#### **Privilege mode Authentication**

You can’t define custom lists, but you can change the authentication enable default list.

## Authorization

### Authorization Lists

Authorization lists are created similarly to authentication lists:

```
R(config)# aaa authorization LIST-TYPE {LIST-NAME|default} METHODS-LIST 
! LIST-TYPE can be:
  commands         For exec (shell) commands.
  configuration    For downloading configurations from AAA server
  exec             For starting an exec (shell).
  network          For network services. (PPP, SLIP, ARAP)
  reverse-access   For reverse access connections
! methods in METHODS-LIST are similar to authorization methods, but there's a new one:
  if-authenticated  Succeed if user has already authenticated.
```

### Applying authorization lists

Authorization usually works hand in hand with. But it can be also used with default privilege levels.

#### **Authorize access to EXEC mode**

By default, the VTY privilege level is set to 1. You can change this level if you set

```
R(config-line)# privilege level LEVEL
```

This means all users that login via the VTY will be assigned this privilege LEVEL.

You can use a dynamic method of assigning users with an authentication level (radius, tacacs, local) or you can fallback to the line configuration (if-authenticated).\
First define the EXEC-LIST-NAME method list, and then you apply it on the vty line:

```
R(config-line)# authorization exec {EXEC-LIST-NAME|default}
```

The same is true for console lines too, except by default authorization always succeeds and assigns the users to level 15 (Applies to router only, not switches), regardless of the authorization exec list defined on the console line. To enable the use of the list, you must also run:

```
R(config)# aaa authorization console
```

#### **Authorize commands**

```
R(config-line)# authorization commands PRIV-LEVEL {COMMANDS-LIST-NAME|default}
```

When command authorization is enabled, the router will check if the user is authorized to run the specified command. It is useless to authorize both exec and commands against the local database since the privilege level defined there will be the same for both authorization types. It makes sense to authorize commands against another server where allowed commands for each user can be defined.\
A user is able to run all commands enabled for its privilege level and all inferior levels, but authorization will only be checked for the commands enabled at the PRIV-LEVEL specified in the authorization command.\
By default, only exec commands are checked. To enable authorization for configuration commands also, use:

```
R(config)# aaa authorization config-commands
```

#### **Authorize PPP**

To enable authorization for PPP, use:

```
R(config-if)# ppp autorization {default|LIST-NAME}
```

## Accounting

Accounting keeps track of the users’s actions while connected to the system.

### Accounting Lists

To define an accounting list, use:

```
R(config)# aaa accounting LIST-TYPE {default|LIST-NAME} {start-stop|stop-only|none} METHOD-LIST
!LIST-TYPE:
 network           For network services. (PPP, SLIP, ARAP)
 connection        For outbound connections. (telnet, rlogin)
 exec              For starting an exec (shell).
 system            For system events.
 commands          For exec (shell) commands.
! METHOD-LIST:
 group radius      Uses RADIUS
 group tacacs+     Uses TACACS+
```

### Applying accounting lists

#### **Line accounting**

To enable line accounting, use:

```
R(config)# accounting {exec|connection|commands|resource} {default|LIST-NAME}
! exec = info about EXEC sessions
! commands = info about commands issued
! connections = info about outbound connections
! resource = info about passed and failed authentications
```

#### **Interface accounting**

#### **PPP Accounting**

To enable PPP accounting, use:

```
R(config-if)# ppp accounting {default|LIST-NAME}
! info about PPP sessions, including packet and byte count
```


# Controlling CLI Access

## CLI Modes

* User EXEC Mode

  ```
  R>
  ! To enter the Privileged Exec Mode
  R> enable
  ```
* Privileged EXEC Mode

  ```
  R#
  ! To go back to the User Exec Mode
  R# disable
  ```
* Configuration Mode

  ```
  ! To enter config mode
  R# config terminal
  ! To exit config mode:
  R(config)# end
  ```

To protect access to the Privileged EXEC mode, use:

```
R(config)# enable {password PASS | secret SECRET}
! secret uses a better encryption algorithm than password encryption
```

### Custom Privilege Levels

These commands are not compatible with [AAA](/security/aaa-101) mode of operation.\
By default, there are only 3 privilege levels:

* Level 0 = no rights
* Level 1 = User EXEC Mode
* Level 15 = Privileged EXEC Mode

You can define new privilege levels and assign commands to them, so they can be ran by lower level users:

```
R(config)# enable level LEVEL {password PASS | secret SECRET}
! set a password for the specific level
R(config)# privilege COMMAND [all] level LEVEL STRING
! All commands starting with STRING will be allowed at the specified level
! all - all suboptions will be allowed at the specified level
! COMMAND - is the parent command of the STRING command.
!        You shoud start building a tree from exec
```

### Role Based CLI Access

This feature is similar to the Custom Privilege levels, but it requires [aaa new-model](/security/aaa-101)\
With this feature you can create custome views with different access to the CLI commands.\
First you need to be in the root view in order to create other views. To move to the root view, use:

```
R# enable view
```

To access the view, you can use the enable secret or password configured on the router. Then, go into the configuration mode and create a view, and specify a secret for it:

```
R(config)# parser view VIEW-NAME
R(config-view)# secret STRING
```

Then, start adding commands to the view, using a similar approach as with Custom Privilege Levels:

```
R(config-view)# commands COMMAND {include|exclude|include-exclusive} [all] STRING
! All commands starting with STRING will be allowed at the specified level
! all - all suboptions will be allowed at the specified level
! include - adds the command to this view
! exclude - removes the command from this view
! include-exclusive - adds the command to this view and removes it from other views
! COMMAND - is the parent command of the STRING command.
!        You shoud start building a tree from exec
```

There’s also the option of configuring a super view. This view can only be a collection of other views:

```
R(config)# parser view SUPER-VIEW superview
R(config-view)# secret SECRET
R(config-view)# view VIEW-NAME
```

You can move from one view to another using the command:

```
R# enable view [VIEW-NAME]
```

an AAA attribute (cli-view-name) can be passed from an AAA server to enable users automatic access to a view.

## CLI Sessions

### Local or Remote CLI

For Local access to the device you must use the Console or AUX port. To configure Local CLI access, use:

```
! Console:
R(config)# line console 0
! AUX:
R(config)# line aux 0
```

For remote CLI acces you must use Telnet or SSH to connect to the device. To configure remote CLI sessions, configure the Virtual Terminal Interfaces using:

```
R(config)# line vty LINE-START [LINE-END]
! Usually, LINE-START=0, LINE-END=4
```

### Protecting line access

By default, the console and AUX ports allow access without asking for credentials. To enable the router to ask for credentials, use the command:

```
R(config-line)# login [local|tacacs]
! no params: use line password (default on vty)
! local: use locally defined users
! tacacs: use tacacs authentication
```

If line password is used but not defined, login will fail. To set the line password, use:

```
R(config-line)# password PASS
```

When using line password, you can also set the default privilege level of authenticated users:

```
R(config-line)# privilege level [0-15]
```

When using locally defined users, you must first set a username:

```
R(config)# username USER [privilege LEVEL] {password PASS | secret SECRET}
! default LEVEL: 1
```

The privilege level of the USER will be used when connecting on the line.

A VTY line can be used for both incoming and outgoing connections. You can define the protocols allowed on each line, using:

```
R(config-line)# transport {input|output} PROTOCOL
! PROTOCL = usually telnet or ssh
! input = incoming connections
! output = outgoing connections
```

Protocols defined with the input keyword are allowed for connections to the terminal line, while protocols defined with the output keyword are protocols that can be used to connect from that line to another host.

## Password Encryption

By default, when a configuration file is saved, passwords are saved in clear text. You can enable automatic encryption of passwords in the config files, using:

```
R(config)# service password-encryption
```

When the configuration files are saved remotely, the passwords are still sent in clear text.\
For better encryption, use secret instead of password when available

## SSH

Before using SSH, a key must be generated, but to generate a key, you need a hostname and domain name.

```
R(config)# hostname HOST
HOST(config)# ip domain-name DOMAIN
HOST(config)# crypto key generate rsa
! You will be asked for the key size.
```

### SSH Server

```
R(config)# ip ssh {timeout SEC|authentication-retries RETRIES}
R(config)# ip ssh version {1|2}
! By default both v1 and v2 users are allowed
! You can chose only one version by selecting it in this command
```

### SSH Client

```
R# ssh -l USER SERVER-IP
```

### SCP Server

[AAA authentication](/security/aaa-101#authentication) must be enabled for SCP. To enable the SCP server, use:

```
R(config)# ip scp server enable
```

The users must pass an aaa login authentication and an aaa exec authorization to use scp.


# Control Plane

## CoPP – Control Plane Policing

Control Plane Policing is used to apply policy maps to traffic going to or coming from the control plane. This feature also mitigates DoS attacks by filtering traffic that arrives at the processor.\
First, you should define a [policy-map using MQC](https://nyquist.eu/qos-101-classifying-and-marking/#2_MQC_Modular_QoS_CLI). Then, apply this policy map to the control-plane:

```
R(config)# control-plane
R(config-cp)# service-policy {input|output} POLICY-MAP
```

To monitor Control Plan Policing use:

```
R# show policy-map control-plane
```

## CoPPr – Control Plane Protection

Control Plane Protection is similar to the policing feature, except it offers a more granular access to the control plane functions.\
You still need to define a [policy-map using MQC](https://nyquist.eu/qos-101-classifying-and-marking/#2_MQC_Modular_QoS_CLI). But now you can apply it on a sub-interface of the virtual control-plane interface:

```
R(config)# control-plane {host|transit|cef-exception}
! host - controls traffic that is directly destined for one of the router interfaces
! transit - controls traffic that is software-switched
! cef-exception - controls traffic that cannot be switched by CEF
R(config-cp-transit)# service-polcy input POLICY-MAP
```

### Layer 4 Port Protection

On the host subinterface you can use a special type of policy-map to deny access-to specific ports.\
First create the special class-map and policy-map:

```
R(config)# class-map type port-filter [match-all|match-any] PORT-FILTER-CLASS
R(config-cmap)# match [not] {closed-ports|port {tcp|udp} START-PORT [END-PORT]}
```

The **closed-ports** keyword should mean all ports that are not open on the control plane. Filtering traffic to the closed ports should spare the processor from unnecessary work.\
To see a list of open ports, use:

```
R# show control-plane host open-ports
```

Beware of the fact that some protocols (usually routing protocols) will not show up with open ports, so they should be specifically allowed if they are used.\
Then create the policy map:

```
R(config)# policy-map type port-filter PORT-FILTER-POLICY
R(config-pmap)# class PORT-FILTER-CLASS
R(config-pmap-c)# drop
```

In the end, apply it on the host subinterface:

```
R(config-cp-host)# service-policy type port-filter input PORT-FILTER-POLICY
```

### Queue Threshold Protection

Similar to the previous feature, you can set how long the queue of the control-plane host subinterface can be. Follow these steps to define the class-map and the policy-map:

```
R(config)#class-map type queue-threshold [match-all|match-any] QUEUE-TH-CLASS
R(config-cmap)# match [not] {protocol [PROTOCOL]|host-protocols}
! host-protocols - any open TCP/UDP port on the router
R(config)# policy-map type queue-threshold QUEUE-TH-POLICY
R(config-pmap)# class QUEUE-TH-CLASS
R(config-pmap-c)# queue-limit LIMIT
```

Then, apply it on the host subinterface:

```
R(config-pmap-host)# service-policy type queue-threshold input QUEUE-TH-POLICY
```

### Management Interfaces

Additionally, on the host subinterface, you can define what management protocols are allowed on each physical interface:

```
R(config-cp-host)# management-interface INTERFACE allow [PROTOCOL]
```

## Control Plane Logging

For traffic that arrives at the control plane (that is traffic that is not cef-switched), the control-plane can limit the amount of logging it does.\
To enable this feature, first define a logging class-map and policy-map:

```
R(config)# class-map type logging match-all LOG-CLASS
! Match by input interface:
R(config-cmap)# match [not] input-interface INTERFACE
! Match on source or destination address
R(config-cmap)# match [not] ipv4 {destination-address|source-address} IP-ADDR
! Match on packet action
R(config-cmap)# match [not] packets {dropped|error|permitted}
R(config)# policy-map type logging LOG-POLICY
R(config-pmap)# class LOG-CLASS
R(config-pmap-c)# log [interval SEC|total-length|ttl]
! interval - logs packets at the specified interval
! total-length - logs also packet length
! ttl - logs also ttl value
```

Apply the logging policy-map to the control plane interface or one of its subinterfaces:

```
R(config)# control-plane [host|transit|cef-exception]
R(config-cp)# service-policy type logging input LOG-POLICY
```


# Switch Security


# Switchport Traffic Control

## Strom Control

The Storm Control feature, will disable the interface as soon as a specific threshold is passed. The threshold is measured every 1 second. The threshold can represent the amount of broadcast, multicast or unicast traffic and it can configured with:

```
! As a percentage of bandwidth:
Sw(config-if)# storm-control {broadcast|multicast|unicast} level LEVEL [LEVEL-LOW]
! As bandwidth in bps:
Sw(config-if)# storm-control {broadcast|multicast|unicast} bps BPS [BPS-LOW]
! As packets per seconds:
Sw(config-if)# storm-control {broadcast|multicast|unicast} pps PPS [PPS-LOW]
```

All traffic on the interface will be blocked when the rising threshold is passed. It will be resumed when traffic falls under the falling threshold.\
L2 multicast traffic used for control (BPDU, CDP) is not affected, but L3 multicast control traffic (routing protocols) is affected by this feature.\
The interface can be shutdown or it can generate a SNMP trap when the threshold is passed:

```
Sw(config-if)# storm-control action {shutdown|trap}
```

The suppression levels can be monitored with:

```
Sw# show storm-control [INTERFACE] [broadcast|multicast|unicasts]
```

## Small Frames

Frames smaller than 67 bytes are not counted by storm-control, but a similar mechanism can be enabeld:

```
! 1. Enable globally:
Sw(config)# errdisable detect cause small-frame
! 2. Enable per interface:
Sw(config-if)# small-violation-rate PPS
```

When the PPS threshold si passed, the port is errdisabled.

## Protected Ports

When 2 ports are defined as protected in a VLAN, they are completly isolated and cannot exchange traffic between them. They can exchange frames with other non-protected ports. It is similar to a private VLAN implementation, but only hardware switched packets are affected. Process switched packets are not affected by this feature.\
To define a proteced port, use:

```
Sw(config-if)# switchport protected
```

## Port Blocking

Port Blocking can be used to disable flooding of multicast, broadcast or unknown unicast from one port to others.

```
Sw(config-if)# switchport block [multicast|unicast]
```


# Switchport Port Security

Port Security restricts the number of stations that are allowed to access a switch port.

## Define allowed hosts

Each time a host attempts to send a frame, the source MAC address is added to the list of secure MACs. This list of secure MAC addresses has a limited size, and it can be configured with these types of secure MAC addresses:

* **Static**: manually configured. Will appear in the running config.

  ```
  Sw(config-if)# switchport port-security mac-address MAC-ADDR
  ```
* **Dynamic**: learned from the source MAC of the frames that enter the port. They are not part of the config and will be lost on reload
* **Sticky**: When enabled, it converts dynamic addresses into static sticky addresses so they will appear in the running config.

  ```
  Sw(config-if)# switchport port-security mac-address sticky
  ```

  . Static sticky addresses can also be added using:

  ```
  Sw(config-if)# switchport port-security mac-address sticky MAC-ADDR
  ```

## Enable port-security

To enable port security, the port must be statically set to access or trunk:

```
Sw(config-if)# switchport mode {access|trunk}
```

To enable port security, use:

```
Sw(config-if)# switchport port-security
```

By default, only 1 mac address will be allowed on this interface. You can modify the number of allowed mac addresses, using:

```
Sw(config-if)# switchport port-security maximum MAX
```

On trunk ports, the maximum value can be set per vlan:

```
Sw(config-if)# switchport port-security maximum MAX vlan [VLAN-LIST]
```

if VLAN-LIST is missing, then a maximum for each vlan is set.

## Security Violation

A security violation occurs when the maximum number of secure MAC addresses have been added to the list and a new station attempts to send a frame, or when an address learned or configured on one secure interface attempts to send a frame on another secure interface in the same VLAN.\
The default security violation mode is “shutdown” but this can be changed with:

```
Sw(config-if)# switchport port-security violation {protect|restrict|shutdown [vlan]}
```

* **protect**: Frames sent by hosts that are not in the secure list are dropped. No notification occurs
* **restrict**: Just like protect, but also sends a SNMP trap, a syslog message is logged and a violation counter increments
* **shutdown**: The interface is errdisabled. The switch sends a SNMP trap, a syslog message is logged and a violation counter increments
* **shutdown vlan**: Only the offending vlan is errdisabled. The switch sends a SNMP trap, a syslog message is logged and a violation counter increments

Recovery from the errdisabled mode can be manually (shut/no shut) or automatically:

```
Sw(config)# errdisable recovery cause psecure-violation 
```

## Clearing allowed hosts list

### Manual

The list of secure MAC addresses can be cleared manually, using:

```
Sw# clear port-security {all|configured|dynamic|sticky} [address MAC-ADDR|interface INTERFACE]
```

### Auto

The addresses in the list of secure MAC addresses can be aged out if you set:

```
Sw(config-if)# switchport port-security aging time SEC
```

The aging can be absolute (the address will be aged out after the configured interval) or relative to inactivity (the address will be aged out after an interval of inactivity equal to the configured value).

```
Sw(config-if)# switchport port-security aging type {absolute|inactivty}
```

Normally, the aging process happens only for dynamic addresses. It cannot be used for sticky addresses, but it can be used for static addresses, using:

```
Sw(config-if)# switchport port-security aging static
```

## Verification

To verify port-security status, you can use:

```
Sw# show port-security [address MAC-ADDR|interface INTEFACE]
```


# DHCP Snooping and DAI

## DHCP Snooping

DHCP snooping can prevent unauthorized DHCP servers to reply to DHCP requests. A switch can define interfaces as trusted or untrusted. A trusted interface is where a DHCP server should be connected. On such interfaces, DHCP server messages are allowed. On all other untrusted ports, DHCP server messages are droped. Also, while this feature runs, the switch builds a DHCP binding database, where it maps all MAC addresses to the IP addresses they received via DHCP.\
To enable DHCP snooping, use:

```
! 1. Enable DHCP Snooping globally
Sw(config)# ip dhcp snooping
! 2. Then enable DHCP snooping on the VLAN
Sw(config)# ip dhcp snooping vlan VLAN-RANGE
! 3. Set up the interfaces as trusted. By default they are untrusted
Sw(config)# interface INTERFACE
Sw(config-if)# ip dhcp snooping trust
```

### Option 82

DHCP Option 82 allows a DHCP server to identify a host by the port on the switch it connects to, in addition to the host MAC Address.\
To enable the switch to add the Option 82 field to the DHCP request message, use:

```
Sw(config)# ip dhcp snooping information option
```

You can also limit the number of DHCP packets that are received on an interface, using:

```
Sw(config)# ip dhcp snooping limit raate PPS
```

To modify the default information that is added, use:

```
Sw(config)# ip dhcp snooping information option format remote-id [string STR|HOSTNAME]
```

By default, when DHCP snooping is on, a switch will drop DHCP requests with Option 82. The switch considers that only it can add this information and it considers those requests as illegitimate.\
This is a good idea on access switches, but on an aggregation switch, it might end up in dropping DHCP requests that had their option 82 inserted by legitimate access switches. To permit such requests, use:

```
Sw(config)# ip dhcp snooping information option allow-untrusted
```

DHCP Snooping is known to add empty giaddr field in the DHCP messages, which will make most DHCP servers ignore them. There are 2 solutions:\
1\. make the server trust DHCP messages with empty giaddr field

```
R(config)# ip dhcp relay information trust-all
```

2\. disable option 82 insertion on the switches

```
Sw(config)# no ip dhcp snooping information option 
```

### DHCP Snooping Binding Database

Information regarding the legitimate hosts on the network are stored in the DHCP Snooping Biding Database. This database contains the MAC address and the IP address that was offered in the DHCP process. This database is lost upon restart. To prevent this, you can specify a DHCP Snooping Binding Database Agent, after which you can set a location where to save this database:

```
Sw(config)# ip dhcp snooping database PROTOCOL://FISER
```

You can also add static information to the database, using:

```
Sw(config)# ip dhcp snooping binding MAC vlan VLAN-ID inteface INTERFACE expiry SEC
```

To verify, use:

```
Sw# show ip dhcp snooping database [detail]
```

## IP Source Guard

You can limit traffic on an untrusted port to a single source IP or MAC address, the ones that are found in the Snooping Binding Database for the specified port. A port ACL is applied to the interface in order to block all other traffic. This ACL will take precedence over any other router ACL or VLAN map that might affect the port.\
To enable IP Source Guard, use one of the following:

```
Sw(config-if)# ip verify source
!Enables IP address filtering
Sw(config-if)# ip verify source port-security
!Enables both IP and MAC address filtering
```

If using the port-security option, the MAC address in the DHCP packet is not learned as a secure address. It will be learned only when it starts to send non-DHCP traffic.\
To verify, use:

```
Sw# show ip source binding [IP-ADDR] [MAC] [dhcp-snooping|static] [interface] INTERFACE [vlan VLAN-ID]
```

If DHCP is not an option, or there are static hosts, you can enable IP Source Guard even for them, using:

```
Sw(config-if)# ip verify source tracking port-security
```

This will track the first non-DHCP MAC it receives and will only let it send and receive packets.

## Dynamic ARP Inspection (DAI)

DAI intercept ARP Replies and will drop those for which the MAC-IP binding is not found in the DHCP Snooping Binding Database. This means that you need to enable DHCP Snooping in order to run Dynamic ARP Inspection.\
To enable DAI, use:

```
Sw(config)# ip arp inspection vlan VLAN-RANGE
```

DAI intercepects ARP replies only on untrusted interfaces. To set an interfaces as trusted, use:

```
Sw(config-if)# ip arp inspection trust
! Usually only host interfaces are set as trusted
```

To verify, use:

```
Sw# show ip arp inspection {interfaces | [statistics] vlan VLAN-RANGE}
```

You can set a limit for the number of ARP requests that you receive on a port:

```
Sw(config)# ip arp inspection limit {rate PPS [burst interval SEC]|none}
!Default: 15 PPS for untrusted interfaces
```

If the rate goes over the limit, the port is errdisabled. To auto enable it, use:

```
Sw(config)# errdisable recovery cause arp-inspection interval INTERVAL
```

### ARP ACLs

When DHCP is not available, you can still enable ARP inspection using ARP ACLs.\
To set up an ARP ACL, first define it:

```
Sw(config)#arp access-lists ACL-NAME
Sw(config-arpacl)# permit ip host SRC-IP mac host SRC-MAC [log]
```

Then apply the ACL on a VLAN:

```
Sw(config)# ip arp inspection filter ACL-NAME vlan VLAN-RANGE[static]
! static - adds an explicit deny to the end of the ARP ACL-NAME
```

### Extra Validations

Apart from discarding ARP packets with invalid IP-to-MAC bindings, DAI can perform additional validation checks on ARP packets:

```
Sw(config)# ip arp inspection validate {[src-mac]|[dst-mac][ip]}
! src-mac: checks the source MAC address in the Ethernet header against the sender MAC address in the ARP body
! dst-mac: check the destination MAC address in the Ethernet header against the target MAC address in ARP body
! ip: checks the ARP body for invalid IP addresses, like 0.0.0.0 or 255.255.255.255
```




---

[Next Page](/llms-full.txt/1)

