Tuesday, January 12, 2016

Quantifying traffic policer rate with Burst-Size value - test method and calculation

Quality Of Service (QoS) dictates how a packet or flow is handled in networking world. Task is not as simple as the sound of this magic three letter word. Many functionalities of QOS work together to accomplishes this trivial task.  QoS complexity is dictated by number of modifications performed to return intended result, QOS  concepts are hard to perceive for many networking professionals.

Let's add one of QoS functionality "POLICING/RATE-LIMITING"

Traffic policing is one of the commonly used QoS feature. Policer/rate-limiter helps to allow only defined packet rate for interested flow. Various other sub-tasks like set queue.no or mark  a specific bit in packet for violated and conformed actions are also possible.

Typical traffic policing configuration looks similar to.
 Police cir 100 mbps bc 200ms pir 200 mbs bc 20 ms conform transmit exceed set-dscp 3 violated drop

above config allows traffic rate  of 100kbs, sets dscp value 3 for flow between 100mbps to 200 mbps and drop any further traffic.

Most networking engineers don't really know that traffic is not actually policed at 100 mbbs, there is more to it, BC helps to define it.

What does "BC - Burst Count" do?
Why my actual policer rate is more than defined rate?
How does it helps on practical traffic flows?

Let's find answer here.

Burst count helps to adjust policer rate to absorb traffic burst. Real world traffic flow is bursty in nature. Handling burst help to have handle on policer rate without increasing the traffic drop point.
 very rare to see a constant rate of data flow, even if you observe high traffic rate, it is definitely constituted by several mice flows than one giant elephant flow.

BC in time representation translates to Bytes based on port speed.

for a 100mbps policer rate  200 ms BC
Policer bandwidth in Bytes is,

(100 x  1000 x 1000 )  bps x 0.2 Sec  = 2500 KBytes
--------------------------------------------
  8 bits in a byte

On a 10gbps port speed this translates  (2500 KB / 10gbps)to, 2msec of burst duration.

Bandwidth rate in 2msec on 100mbps link becomes 0.2mbps. Effective policer rate is 100.2 Mbps


Typical BC metrics are either in time (ms/micro-sec/sec) or in Bytes. Now that you know the conversion. effective policer rate can be easily determined. 

Friday, January 1, 2016

Data Center Switch Market - Black magic in White Box

Deploy, Manage and operate Data center switches similar to the way you operate a server. Buy commodity hardware and run operating system you like. These two ideas gave birth to WHITE BOX switches to Switching market.

In campus networking and Data Center  networking Top of Rack (TOR) or Leaf switches are the most deployed networking infrastructure. Every network interconnect should go through LAN switches, In Data Center networking TOR is the 1st server interconnect point. Every server gets a link to a TOR switch through direct or indirect extended links. White box evolution is promises to make a big impact on TOR market with low priced commodity switches. Data Center switching has made a giant stride in recent time, Switches are specially built instead of using general-purpose switches. There has been a quantum jump in amount of traffic handled by data center switches. I have used 100Mbps port-speed switches to connect server, Gone or the days!

White box echo system consists of,

  • Merchant silicon companies - Networking ASICS - Ex: BRCM, xpliant, NxP, Intel
  • Bare Metal Switch providers -  Ex: Quanta computer Inc, Pic8, Accton, Celestica
  • Network OS Ex: Cumulus, Pic8, Big Switch Networks, Juniper, Dell etc

ASIC, HW and OS together makes a switch. Choice for each one of these through various vendors makes life easy for all customers. This Healthy competition is laying founding stones for the fruitful future of White box switches. 

I am neither a support nor an opposer to White box solutions. My data center domain expertise only puzzles me with following questions. 

  1. Server OS != Switch OS. I couldn't accept this point.
    • x86 architecture has been there for a while. hypothetically Ever since computer industry has evolved into mainstream.
    • Server OS is a Very very big market to keep different providers busy.Desktop, Application hosting environment, Cloud, Campus, cellphone and lab)
    • Networking processors have to go through multiple sprints to get the maturity equivalent to PC processors. Evolving protocols needs newer capabilities in ASIC.
  2. Support Onus - Asic, Switch Manufacturer and OS, Out of these three pillars who will take ownership for any issues.
  3. Catch-Up with standard - Networking giants like Cisco, Juniper and Brocade have edge over Merchant silicon vendors in many new protocols.
  4. Support cost is directly proportionate to  knowledge base. New Box/OS - means new training cycle for IT engineers
Certainly WHITE BOX solutions is an important catalyst for SDN and NFV evolution. Data Center and enterprise networking is going through a big consolidation phase. White Box battle will certainly make a big hole in existing networking vendors revenue. I am sure they will have plans to sail through this head wind. 

Open standards is must to increase innovation and hence tackle digital divide across the globe. Only time can reveal the effectiveness of this black magic. 




    Sunday, December 13, 2015

    Going away from silos to Hyper Converged Systems

    A typical starting statement  of a IT support help desk has statements like, 'xyz' in app did not work, getting unexpected error. 

    Firstly, support executive tries to isolate problem domain. Upon making 1st level triage, reassigns ticket to domain expert. Domain expert looks at the issue and expresses his option. More often load handling of infra is blamed. IT team takes it further to scale up specific infra. Depending on the app importance immediate or delayed budget gets this new HW. IT provisioning team integrates new node and issue is finally resolved.

    Though the solution flow looks simple, often it takes weeks or months to add a new compute resource in traditional IT handling methods. Some trivial tasks involved in adding new infra are: Data Center rack space, physical network connection, provisioning, OS installation and app migration. 

    Yes, gone are the days. Virtualization converts these trivial tasks into few simple click events. 

    • A simple clone operation and network param change creates new node. 
    • Scaling up compute, storage and memory can happen in few clicks.

    having said that, still adding a new node to existing cluster can take long cycles. This can add up lead time. 
    • In NAS and SAN environment scaling up beyond current storage limits take lot of time. 
    • CPU cycle still follow traditional delays, It's difficult to add pur compute infra by itself.
    • HA and recovery from fault is time consuming

    End of problem statement leads to solution segment. Here you go,
    Many startups are working on solutions to remove above mentioned silos. End goal is to reduce lead time involved in compute, storage and memory expand and shrink operation.  companies to watch out in this space are, 
    1. Nutanix - San Jose based startup, offers different models of Hyper converged systems. Works with all leading Hypervisor vendors. recently came-up with its own hypervisor as well.
    2. Nimboxx - built on the same platform as existing Nimboxx solutions, combining compute, storage, networking and security into a single system. High availability comes from the Nimboxx MeshOS operating environment
    3. Pivot3 The vSTAC OS dynamically aggregates, load-balances and optimizes shared storage, compute and network resources within the Pivot3 
    4. Simpivity
    Other than start-ups few established players have their own solutions as well.

     Riverbed, Supermicro.

    Hyper converged systems will be a very interesting space to watchout in 2015 and 2016. Software defined Storage (SDS) is expected to take good shape with hyper converged systems. 

    Friday, August 15, 2014

    Wireless Networking challenges

    I must say Data (Voice+packet) that can be carried per Hz (Unit for Frequency) rapidly gone up in this decade. Thanks to intelligent brains in finding better results in Signal processing and Mobile communication. I am sure we will witness much faster mobile communication and intelligent nodes (Your AC, Smoke detectors, Window, lighting etc) in coming days. Wireless networking is definitely a happening field in networking sector. Wireless radios have matured from  proprietary technologies to heavily into standard based solution in a very short span of time.

    Primary challenges in front of wireless networking startups,

    Improve SNR - Signal to Noise ration

    1. Increase Signal penetration (Gigantic Base Station to Pico cells). LOS (Line of Sight) performance in Non-LOS conditions is the goal.
    2. Intelligence to Antenna (Reducing transmission loss, noise cancellation and coverage improvement)
    3. Beat Shannon's law

    From service providers point of view:

    1. Mobility - Signal handoff, LTE advanced and WiMax advanced made a breakthrough here.
    2. Wireless device management - Software Defined radio
    3. Monitoring - Intelligent Load sharing
    There may be numerous other challenges but above list offers a simple baseline.

    Interesting part of this post is to cover wireless networking startups but wait for my next post to know about Kum networks, DIDO

    Debugging with GDB - Lets decypher Segmentation Fault

    I want to be a philanthropist. Hold on!  Are you looking for any link to get money? Get out of my website if you have wrongly landed here in search of money. I am sure Google (Search engine is not the key word any more) does a good job in listing search results. 

    I only offer information in netglutton, after all information is wealth. I recently gathered more information about a tool i use almost every day in my Software Engineer role. a magic three letter tool "GDB - GNU Debugging".

    Process crash or Core singles out as "Critical" defect among tons of defects in any software system. It is so severe that both test engineer and developer have hard time in resolving process crash. As a software test engineer i have encountered 100s of Process Crash scenarios. 

    Why it happens?
    Simply, unhandled exception. Yes, Process crash happens due to unhandled exceptions like Null Pointer, Memory error, EOF etc in code. As a test engineer i have not only seen simple way to trigger a crash but also a most difficult ones. 

    DO NOT TRY THIS: 
    If you have root access, find a non-restartable process (ps -ef) and do "kill -9 or -11".

    "Breaking Point" is very critical issues in code. I really dont want to get into the reason for Crash/Cores. A C/C++ developer must know about GNU Debugging (GDB). Software test engineer should also know about collecting core file, attaching to GDB and finding BT (BackTrace short form).

    Unhandled exception OR Null pointer OR Segmentation Fault in process leads to process restart. Process failure leads to core file. Back trace using GDB helps to find breaking function call /memory location/ value. 

    With above contextual information, go through following help links for more information about How to GDB?



    Sunday, July 27, 2014

    Midonet - Network Virtualization Solution from Midokura:

    SDN is not a story anymore, Several players have solutions to try. With this blog post i am going to share my read/research about Midonet -  A Network Vitalization Solution from a start-up "Midokura". Unlike other leading SDN providers, Midokura prime focus is fixed at Decentralization.

    Instead of a designated controller based approach, Midonet has taken a Distributed controller approach. Every Hypervisor will act as a virtual Router hence highly Distributed. Distributed routing intelligence combined with Border Gateway (Physical Router) controls network traffic to/from datacenter from/to Internet.

    Midonet operation explained in few words,

    Central flow DB either gives information about destination node and flow or finds a path to destination. Source establishes a GRE tunnel to destination. Destination could be another Hypervisor or Gateway connected to Internet. Packet intercepted in Hypervisor at kernel space, encapsulated inside a GRE.

    Key Elements involved in Midonet solution:

    Hypervisor interconnect: Midonet simply expects a ip switching/routing reachability between all hypervisors and Gateway. No vendor dependency.

    Agent: Every Hypervisor needs Midonet Agent installation. Agent derives flow information from central DB for 1st packet, rest of the packets to same destination will directly go through established tunnel to destination Hypervisor.

    Gateway: x86 server with Midonet Agent. Talks to external network in E-BGP.

    Central Network flow DataBase: All Midonet agents subscribe to this DB. DB contains every information about every VM.

    Midonet API, GUI, Orchestration: API offers programmable interface to View/control Agents. GUI does the same to graphical user. Easy to integrate with cloud orchestration tools like OpenStack and CloudStack.

    Now, please read Midonet operation explanation once again with Midonet elements in mind.

    A VirtualMachine wants to reach a destination, of course VM is inside a hypervisor. Hypervisor gets the packet to send. Midonet Agent intercepts the packet, finds tunnel information from Network DB. Establishes a tunnel to destination. Tunnel destination could be another Hypervisor or Gateway (for external traffic).

    Hope  you have fill picture of Midonet's SDN offering.

    Sunday, July 13, 2014

    First Thoughts on OpenContrail

    Thanks to Juniper Networks and Bay Area Network Virtualization Meet-up group for a wonderful OpenContrail hands on session. Juniper product marketing folks created an awesome environment to play and learn their solution. I appreciate their openness, indeed a rare quality among big companies. Especially from the one, Juniper, which is under lot of pressure to perform in changing DataCenter networking field.

    What does OpenContrail Do?

    Open, standard based  solution to do Network virtualization and service automation for cloud network.
    It has 3 important stakeholders,
    1. Controller (configuration, control and Analytics
    2. VRouter (Every compute element has VRouter)
    3. Gateway (Exit point to external network) - any MPLS-VPN OR VxLAN and Tunneling supported Router

    Solution expects IP reachability between all nodes very simple flat network is sufficient.
    You can create a IP Pool, Create VM and set its NIC. Whole set becomes a VPC (Virtual Private Cloud).

    Now, Define Network Functions .i.e policies from GUI like access to external IP, IP NAT, Firewall and LB (Load balancing) policies. Without contrail imagine provisioning public/private cloud in DC and making any small changes. If you have ever dealt with network provisioning team you know the pain and delay. Software Defined DataCenter aimed to remove those provisioning and monitoring hurdles, contrail is moving ahead with its solution.

    OpenContrail is bringing so much value to SDN echo system. Solution is completely open. Network provisioning, management and analytics comes easy to scale your DC without much hurdle.

    OpenContrail has so many things to talk about. I would like to talk more in newer posts. For now i would like to complete this post with few question.

    Flat IP connectivity is a big piece to manage. How network bottle neck can be solved?
    How I/O can be managed? Storage is a critical piece in VM provisioning
    How deep analytics can penetrate? How Infrastructure can be handled better with application intelligence?