Showing posts with label Infiniband. Show all posts
Showing posts with label Infiniband. Show all posts

Friday, December 2, 2016

Error polling HP CQ with status WORK REQUEST FLUSHED ERROR status on LSF Platform

I was encountering "Error polling HP CQ with status WORK REQUEST FLUSHED ERROR status" during OpenMPI run and it was occuring randomly.

I suspect it is due to nodes issue. I checked the LSF /opt/lsf/log/sbatchd.log.comp001. It is definitely an authentication issue with AD. I'm using centrify.

acctMapTo: No valid user name found for job 149044, userName(mr_x) failed:Success
runEexec: getOSUid_() failed. Bad user ID


I did a
$ badmin hclose comp001

and then restart centrify services. Alternatively, you can reboot if you want a clean start.

The OpenMPI could run again.

Wednesday, November 4, 2015

Saturday, April 4, 2015

Using ibdiagnet to generate topology of the network.

You can use the ibdiagnet to generate the topology of the IB Network simply by using the "-w" switch
# ibdiagnet -w /var/tmp/ibdiagnet2/topology.top
.....
.....
-I- ibdiagnet database file   : /var/tmp/ibdiagnet2/ibdiagnet2.db_csv
-I- LST file                  : /var/tmp/ibdiagnet2/ibdiagnet2.lst
-I- Topology file             : /var/tmp/ibdiagnet2/topology.top
-I- Subnet Manager file       : /var/tmp/ibdiagnet2/ibdiagnet2.sm
-I- Ports Counters file       : /var/tmp/ibdiagnet2/ibdiagnet2.pm
-I- Nodes Information file    : /var/tmp/ibdiagnet2/ibdiagnet2.nodes_info
-I- Partition keys file       : /var/tmp/ibdiagnet2/ibdiagnet2.pkey
-I- Alias guids file          : /var/tmp/ibdiagnet2/ibdiagnet2.aguid

# vim /var/tmp/ibdiagnet2/topology.top

# This topology file was automatically generated by IBDM

SX6036G Left-Leaf-SW03
U1/P1 -4x-14G-> HCA_1 mtlacad05 U1/P1
U1/P17 -4x-14G-> SX6012 Right-Spine-SW02 U1/P2
U1/P18 -4x-14G-> SX6012 Left-Spine-SW01 U1/P2
U1/P2 -4x-14G-> HCA_1 mtlacad07 U1/P1
U1/P3 -4x-14G-> HCA_1 mtlacad03 U1/P1
U1/P4 -4x-14G-> HCA_1 mtlacad04 U1/P1
U1/P6 -4x-14G-> HCA_1 mtlacad06 U1/P1
.....
.....

Monday, March 30, 2015

Using ibdev2netdev to quickly identify ports

ibdev2netdev is a nice tool to quickly identify ports to ib0

[root@headnode-h99 ~]# ibdev2netdev
mlx4_0 port 1 ==> ib0 (Up)
mlx4_0 port 2 ==> ib1 (Down)

Tools for Performance Test for IB

ibportstate
  • Enables the querying of the logical link and physical por tstates of an IB Port.
  • Displays information such as LinkSpeed, LinkWidth and extended link speed
  • Allows adjusting of link speed that is enabled on any IB Port
# ibportstate LID PortNumber
# Port info: Lid 15 port 1
LinkState:.......................Active
PhysLinkState:...................LinkUp
Lid:.............................15
SMLid:...........................1
LMC:.............................0
LinkWidthSupported:..............1X or 4X
LinkWidthEnabled:................1X or 4X
LinkWidthActive:.................4X
LinkSpeedSupported:..............2.5 Gbps or 5.0 Gbps or 10.0 Gbps
LinkSpeedEnabled:................2.5 Gbps or 5.0 Gbps or 10.0 Gbps
LinkSpeedActive:.................10.0 Gbps
LinkSpeedExtSupported:...........14.0625 Gbps
LinkSpeedExtEnabled:.............14.0625 Gbps
LinkSpeedExtActive:..............14.0625 Gbps
Mkey:............................<not displayed>
MkeyLeasePeriod:.................0
ProtectBits:.....................0
# MLNX ext Port info: Lid 15 port 1
StateChangeEnable:...............0x00
LinkSpeedSupported:..............0x01
LinkSpeedEnabled:................0x01
LinkSpeedActive:.................0x00

Sunday, August 17, 2014

Using iblinkinfo to report link infomation for all links in the fabric

The command iblinkinfo is a useful command to give a good overview of link information for all links in the fabric

[root@node-h00 ~]# iblinkinfo
CA: node-c27 HCA-1:
      0xxxxxxxxxxxxxxx     24    1[  ] ==( 4X          10.0 Gbps Active/ ..... 
CA: node-c26 HCA-1:
      0xyyyyyyyyyyyyyyy    22    1[  ] ==( 4X          10.0 Gbps Active/ ..... 
CA: node-c25 HCA-1:
      0xzzzzzzzzzzzzzzz     29    1[  ] ==( 4X          10.0 Gbps Active/ .....
CA: node-c24 HCA-1:
      0xaaaaaaaaaaaaaaa     28    1[  ] ==( 4X          10.0 Gbps Active/ .....
CA: node-c25 HCA-1:
      0xsssssssssssssss     27    1[  ] ==( 4X          10.0 Gbps Active/ .....
...................
...................
...................

Step 2: Print all information on each node on one single line
[root@node-h00 ~]# iblinkinfo -l 

0xaaaaaaaaaaaaaaaa "          node-c26 HCA-1"     22    1[  ] ==( 4X          10.0 Gbps Active/  LinkUp)==>  0xswitchfa     19   29[  ] "IBM HSSM" ( )
0xvvvvvvvvvvvvvvvv "          node-c25 HCA-1"     29    1[  ] ==( 4X          10.0 Gbps Active/  LinkUp)==>  0xswitchfa     19   28[  ] "IBM HSSM" ( )
0xbbbbbbbbbbbbbbbb "          node-c24 HCA-1"     28    1[  ] ==( 4X          10.0 Gbps Active/  LinkUp)==>  0xswitchfa     19   27[  ] "IBM HSSM" ( )
0xcccccccccccccccc "          node-c23 HCA-1"     27    1[  ] ==( 4X          10.0 Gbps Active/  LinkUp)==>  0xswitchfa     19   26[  ] "IBM HSSM" ( )
0xdddddddddddddddd "          node-c20 HCA-1"     26    1[  ] ==( 4X          10.0 Gbps Active/  LinkUp)==>  0xswitchfa     19   25[  ] "IBM HSSM" ( )
0xeeeeeeeeeeeeeeee "          node-c21 HCA-1"     23    1[  ] ==( 4X          10.0 Gbps Active/  LinkUp)==>  0xswitchfa     19   24[  ] "IBM HSSM" ( )

.....
.....
.....


Step 3: List Down Ports in the Fabric
[root@strawberry-h00 ~]# iblinkinfo -d

Switch: swswswswswswss IBM HSSM:
          19   16[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          19   19[  ] ==(                Down/ Polling)==>             [  ] "" ( )
          19   23[  ] ==(                Down/ Polling)==>             [  ] "" ( )
          19   31[  ] ==(                Down/ Polling)==>             [  ] "" ( )
          19   32[  ] ==(                Down/ Polling)==>             [  ] "" ( )
          19   33[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          19   34[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          19   35[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          19   36[  ] ==(                Down/Disabled)==>             [  ] "" ( )
Switch: swswswswswswswsw IBM HSSM:
          16   31[  ] ==(                Down/ Polling)==>             [  ] "" ( )
          16   32[  ] ==(                Down/ Polling)==>             [  ] "" ( )
          16   33[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          16   34[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          16   35[  ] ==(                Down/Disabled)==>             [  ] "" ( )
          16   36[  ] ==(                Down/Disabled)==>             [  ] "" ( )


For all information:
  1. iblinkinfo(8) - Linux man page  

Tuesday, May 13, 2014

A relook at libibverbs: Warning: RLIMIT_MEMLOCK is 32768 bytes. This will severely limit memory registrations.

There was an prior blog entry written in Oct 2009.libibverbs: Warning: RLIMIT_MEMLOCK is 32768 bytes. This will severely limit memory registrations.

I would like to add on to this entry. In the FAQ 17 from OpenMPI,  17. I'm still getting errors about "error registering openib memory"; what do I do?, the FAQ mentioned about the scheduler

 Make sure that the resource manager daemons are started with unlimited memlock limits (which may involve editing the resource manager daemon startup script, or some other system-wide location that allows the resource manager daemon to get an unlimited limit of locked memory). Otherwise, jobs that are started under that resource manager will get the default locked memory limits, which are far too small for Open MPI.

The files in limits.d (or the limits.conf file) does not usually apply to resource daemons! The limits.s files usually only applies to rsh or ssh-based logins. Hence, daemons usually inherit the system default of maximum 32k of locked memory (which then gets passed down to the MPI processes that they start). To increase this limit, you typically need to modify daemons' startup scripts to increase the limit before they drop root privliedges.

Some resource managers can limit the amount of locked memory that is made available to jobs. For example, SLURM has some fine-grained controls that allow locked memory for only SLURM jobs (i.e., the system's default is low memory lock limits, but SLURM jobs can get high memory lock limits). See these FAQ items on the SLURM web site for more details: propagating limits and using PAM.


Other related Issues

1. For Torque, you may want to tweak the /etc/init.d/pbs_mom See blog entry Default ulimit setting in torque overide ulimit setting

# service pbs_mom restart

2. See also Encountering Segmentation Fault, Bus Error or No output . In that blog, you have to edit /etc/security/limit.conf

* soft memlock unlimited
* hard memlock unlimited

3. If you still have memory issues and using Mellanox IB Cards, do take a look at Registering sufficent memory for OpenIB when using Mellanox HCA

Wednesday, October 9, 2013

Error during installing Mellanox MLNX_OFED 1.5.3 for CentOS 6.3.

I downloaded the Mellanox MLNX_OFED Drivers for RH 6.3 and CentOS 6.3

[root@node-c01 iso]# ./mlnxofedinstall
This program will install the MLNX_OFED_LINUX package on your machine.
Note that all other Mellanox, OEM, OFED, or Distribution IB packages will be r                                                                                                                       ved.
Do you want to continue?[y/N]:y


rpm -e --allmatches --nodeps kernel-mft kernel-mft knem kernel-mft

Starting MLNX_OFED_LINUX-1.5.3-4.0.42 installation ...

Installing kernel-mft RPM
Preparing...                ##################################################
kernel-mft                  ##################################################
Installing knem RPM
Preparing...                ##################################################
knem                        ##################################################
libibumad was not created

This look like the libibumad packages are missing so I did a

# yum install  libibumad* 

But when I launched the ./mlnxofedinstall, the error appear again. Somehow the scripts remove the libibumad packages that I installed and the error occured again.

To resolve the issue, I used MLNX_OFED 2.0.3 and everything was ok.



Friday, August 30, 2013

Diagnostic Tools to diagnose Infiniband Fabric Information

There are a few diagnostic tools to diagnose Infiniband Fabric Information. Use man for the parameters for the
  1. ibnodes - (Show Infiniband nodes in topology)
  2. ibhosts - (Show InfiniBand host nodes in topology)
  3. ibswitches- (Show InfiniBand switch nodes in topology)
  4. ibnetdiscover - (Discover InfiniBand topology)
  5. ibchecknet - (Validate IB subnet and report errors)
  6. ibdiag (Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices)
  7. perfquery (find errors on a particular or number of HCA's and switch ports)
For more information, do look at Diagnostic Tools to diagnose Infiniband Fabric Information

Sunday, January 27, 2013

Diagnostic Tools to diagnose Infiniband Device

There are a few Diagnostic Tools to diagnose Infiniband Devices.
  1. ibv_devinfo (Query RDMA devices)
  2. ibstat (Query basic status of InfiniBand device(s))
  3. ibstatus (Query basic status of InfiniBand device(s))

ibv_devinfo (Query RDMA devices) 
Print  information about RDMA devices available for use from userspace.
# ibv_devinfo

hca_id: mlx4_0
        transport:                      InfiniBand (0)
        fw_ver:                         2.10.2322
        node_guid:                      0002:c903:0045:1280
        sys_image_guid:                 0002:c903:0045:1283
        vendor_id:                      0x02c9
        vendor_part_id:                 4099
        hw_ver:                         0x0
        board_id:                       IBM0FD0140019
        phys_port_cnt:                  2
                port:   1
                        state:                  PORT_ACTIVE (4)
                        max_mtu:                2048 (4)
                        active_mtu:             2048 (4)
                        sm_lid:                 1
                        port_lid:               1
                        port_lmc:               0x00
                        link_layer:             IB

                port:   2
                        state:                  PORT_DOWN (1)
                        max_mtu:                2048 (4)
                        active_mtu:             2048 (4)
                        sm_lid:                 0
                        port_lid:               0
                        port_lmc:               0x00
                        link_layer:             IB

ibstat (Query basic status of InfiniBand device(s))

ibstat is a binary which displays basic information obtained  from  the local  IB  driver.  Output  includes LID, SMLID, port state, link width active, and port physical state.

It is similar to the ibstatus  utility  but  implemented  as  a  binary rather  than a script. It has options to list CAs and/or ports and displays more information than ibstatus.

# ibstat

CA 'mlx4_0'
        CA type: MT4099
        Number of ports: 2
        Firmware version: 2.10.2322
        Hardware version: 0
        Node GUID: 0x0002c90300451280
        System image GUID: 0x0002c90300451283
        Port 1:
                State: Active
                Physical state: LinkUp
                Rate: 40
                Base lid: 1
                LMC: 0
                SM lid: 1
                Capability mask: 0x0251486a
                Port GUID: 0x0002c90300451281
                Link layer: InfiniBand
        Port 2:
                State: Down
                Physical state: Polling
                Rate: 40
                Base lid: 0
                LMC: 0
                SM lid: 0
                Capability mask: 0x02514868
                Port GUID: 0x0002c90300451282
                Link layer: InfiniBand


ibstatus - (Query basic status of InfiniBand device(s))

ibstatus is a script which displays basic information obtained from the local IB driver. Output includes LID, SMLID,  port  state,  link  width active, and port physical state.

# ibstatus

Infiniband device 'mlx4_0' port 1 status:
        default gid:     fe80:0000:0000:0000:0002:c903:0045:1281
        base lid:        0x1
        sm lid:          0x1
        state:           4: ACTIVE
        phys state:      5: LinkUp
        rate:            40 Gb/sec (4X QDR)
        link_layer:      InfiniBand

Infiniband device 'mlx4_0' port 2 status:
        default gid:     fe80:0000:0000:0000:0002:c903:0045:1282
        base lid:        0x0
        sm lid:          0x0
        state:           1: DOWN
        phys state:      2: Polling
        rate:            40 Gb/sec (4X QDR)
        link_layer:      InfiniBand

Thursday, January 17, 2013

OFED Performance Micro-Benchmark Latency Test

Open Fabrics Enterprise Distribution (OFED) has provided simple performance micro-benchmark has provided a collection of tests written over uverbs. The notes taken from OFED Performance Tests README
  1. The benchmark uses the CPU cycle counter to get time stamps without a context switch.
  2. The benchmark measures round-trip time but reports half of that as one-way latency. This means that it may not be sufficiently accurate for asymmetrical configurations.
  3. Min/Median/Max results are reported.
    The Median (vs average) is less sensitive to extreme scores.
    Typically, the Max value is the first value measured Some CPU architectures
  4. Larger samples only help marginally. The default (1000) is very satisfactory.   Note that an array of cycles_t (typically an unsigned long) is allocated once to collect samples and again to store the difference between them.   Really big sample sizes (e.g., 1 million) might expose other problems with the program.
On the Server Side
# ib_write_lat -a

On the Client Side
# ib_write_lat -a Server_IP_address

For more information, do take a look at OFED Performance Micro-Benchmark Latency Test

Tuesday, January 15, 2013

Understanding the Infiniband Subnet

Intel has published a easy-to-understand article on Infiniband Subnetting. "Understanding the InfiniBand Subnet Manager"

From the article:

The InfiniBand subnet manager (OpenSM) assigns Local IDentifiers (LIDs) to each port connected to the InfiniBand fabric, and develops a routing table based off of the assigned LIDs. 
....
....
A typical InfiniBand installation using the OFED package will run the OpenSM subnet manager at system start up after the OpenIB drivers are loaded. This automatic OpenSM is resident in memory, and sweeps the InfiniBand fabric approximately every 5 seconds for new InfiniBand adapters to add to the subnet routing tables. This usage will be sufficient for most installations, and can be controlled using the following commands:

/etc/init.d/opensmd start
/etc/init.d/opensmd stop
/etc/init.d/opensmd restart
/etc/init.d/opensmd status


For more information, read the article.

Monday, January 7, 2013

Errors when running doing ib testing with ib_write_lat

I was doing a ib test using  perftest package which has simple tests for benchmarking IB bandwidth and latency. the 2 simple packages are ib_write_bw and ib_write_lat. 

On the Server side, I launched
# ib_write_lat -a

------------------------------------------------------------------
                    RDMA_Write Latency Test
 Number of qps   : 1
 Connection type : RC
 Mtu             : 2048B
 Link type       : IB
 Max inline data : 400B
 rdma_cm QPs     : OFF
 Data ex. method : Ethernet
------------------------------------------------------------------
 local address: LID 0x03 QPN 0x0065 PSN 0xba11f4 RKey 0x003900 VAddr 0x002ab9bad9600

On the Client Side,
# ib_write_lat -a 192.168.5.1

The errors are as followed
Conflicting CPU frequency values detected: 1200.000000 != 2501.000000
 2       1000          inf            inf          inf
Conflicting CPU frequency values detected: 1200.000000 != 2501.000000
 4       1000          inf            inf          inf
Conflicting CPU frequency values detected: 1200.000000 != 2501.000000
 8       1000          inf            inf          inf
Conflicting CPU frequency values detected: 1200.000000 != 2501.000000
 16      1000          inf            inf          inf
Conflicting CPU frequency values detected: 1200.000000 != 2501.000000
 32      1000          inf            inf          inf

To solve the issues, use "-F" option while running the tests. The flag will ignore "Conflicting CPU frequency" errors. Although there will still be error messages, but with "-F", you will also see the results at least

A better solution is to disable the cpuspeed if you are on CentOS. For more information see blog entry

Wednesday, December 26, 2012

Switching between Ethernet and Infiniband using Virtual Protocol Interconnect (VPI)

This short writeup is a summary of the article Switching between Ethernet and Infiniband using Virtual Protocol Interconnect (VPI). Of course you will need to use the QSA Adapter (QSFP+ to SFP+ adapter) which is the world's first solution for the QSFP to SFP+ conversion challenge for 40GB/Infiniband to 10G/1G. For more information, see Quad to Serial Small Form Factor Pluggable (QSA) Adapter to allow for the hardware


For the full article, see Switching between Ethernet and Infiniband using Virtual Protocol Interconnect (VPI)

Overview
mlx4 is the low level driver implementation for the ConnectX adapters designed by Mellanox Technologies. The ConnectX can operate as an InfiniBand adapter, as an Ethernet NIC, or as a Fibre Channel HBA. The driver in OFED 1.4 supports Infiniband and Ethernet NIC configurations. To accommodate the supported configurations, the driver is split into three modules:
  1. mlx4_core
    Handles low-level functions like device initialization and firmware commands processing. Also controls resource allocation so that the InfiniBand and Ethernet functions can share the device without interfering with each other.
  2. mlx4_ib
    Handles InfiniBand-specific functions and plugs into the InfiniBand midlayer
  3. mlx4_en
    A new 10G driver named mlx4_en was added to drivers/net/mlx4. It handles Ethernet specific functions and plugs into the netdev mid-layer.
Using Virtual Protocol Interconnect (VPI) to switch between Ethernet and Infiniband
Loading Drivers
  1. The VPI driver is a combination of the Mellanox ConnectX HCA Ethernet and Infiniband drivers. It supplies the user with the ability to run Infiniband and Ethernet protocols on the same HCA.
  2. Check the MLX4 Driver is loaded, ensure that the
    # vim /etc/infiniband/openib.conf
    # Load MLX4_EN module
    MLX4_EN_LOAD=yes
  3. If the MLX4_EN_LOAD=no, the Ethernet Driver can be loaded by running
    # /sbin/modprobe mlx4_en
Port Management / Driver Switching
  1. Show Port Configuration
    # /sbin/connectx_port_config -s
    --------------------------------
    Port configuration for PCI device: 0000:16:00.0 is:
    eth
    eth
    --------------------------------
  2. Looking at saved configuration
    # vim /etc/infiniband/connectx.conf
  3. Switching between Ethernet and Infiniband
    # /sbin/connectx_port_config
  4. Configuration supported by VPI
    - The following configurations are supported by VPI:
     Port1 = eth   Port2 = eth
     Port1 = ib    Port2 = ib
     Port1 = auto  Port2 = auto
     Port1 = ib    Port2 = eth
     Port1 = ib    Port2 = auto
     Port1 = auto  Port2 = eth
    
      Note: the following options are not supported:
     Port1 = eth   Port2 = ib
     Port1 = eth   Port2 = auto
     Port1 = auto  Port2 = ib
For more information, see
  1. ConnectX -3 VPI Single and Dual QSFP+ Port Adapter Card User Manual (pdf)
  2. Open Fabrics Enterprise Distribution (OFED) ConnectX driver (mlx4) in OFED 1.4 Release Notes

Friday, December 21, 2012

Quad to Serial Small Form Factor Pluggable (QSA) Adapter


Quad to Serial Small Form Factor Pluggable (QSA) Adapter designed by Mellanox Technologies is the world’s first solution for the QSFP to SFP+ conversion challenge.

The QSA enables smooth, cost-effective, connections between Virtual Protocol Interconnect® (VPI) or 40 Gigabit Ethernet adapters using contemporary QSFP ports and 1 or 10 Gigabit Ethernet networks using existing SFP or SFP+ based cabling. Similarly Ethernet switches with 40Gb/s QSFP ports can connect to servers with 10Gb/s Ethernet NIC ports using QSA.

For more information, see Quad to Serial Small Form Factor Pluggable (QSA) Adapter from Mellanox

Tuesday, October 9, 2012

iWARP, RDMA and TOE

I have captured some basic information on iWARP, RDMA, TOE and RDMA communication....

Remote Direct Access Memory Access (RDMA) allows data to be transferred over a network from the memory of one computer to the memory of another computer without CPU intervention. There are 2 types of RDMA hardware: Infiniband and RDMA over IP (iWARP). OpenFabrics Enterprise Distribution (OFED) stack provides common interface to both types of RDMA hardware.

For more information: iWARP, RDMA and TOE by Linux Cluster

Monday, July 30, 2012

Intel Infiniband Solution - TrueScale Infiniband


Information on Intel® TrueScale InfiniBand Solutions.
You would have probably known Intel buy-over of Qlogic. Not much detailed information are revealed yet on Intel site, but you can find similar information at Qlogic site, but without Intel specific part number

  1. Intel® TrueScale InfiniBand Edge and Director Switches
  2. Intel® TrueScale InfiniBand Fabric Management and Software Tools 
  3. Intel® TrueScale InfiniBand Host Adapters
  4. Intel® InfiniBand Cables

Friday, August 26, 2011

Infiniband HOWTo

Stumble on this short but wonderful tutorial on Infiniband HOWTo by Guy Coates. Although the article is for Debian, but you can apply to CentPS

Sunday, August 7, 2011

Infiniband versus Ethernet myths and misconceptions

I have written my thoughts on the Infiniband versus Ethernet myths and misconceptions. For more information, see Infiniband versus Ethernet myths and misconceptions from Linux Cluster Blog

Only 3 critical myths are discussed.
  1. Opinion 1: Infiniband is lower latency than Ethernet
  2. Opinion 2: QDR‐IB has higher bandwidth than 10GbE
  3. Opinion 3: IB Switch scale better than 10GbE


     Most of my materials are taken from Chelsio White Paper - Eight myths about InfiniBand WP 09-10

    Wednesday, June 22, 2011

    Encountering a Infiniband disconnection to SCSI RDMA Protocol (SRP) target

    Today I encounter this strange error. At /var/log/messages "kernel: ib_srp: ASYNC event= 17 on device= mlx4_0".
    I'm still not sure of it though

    The IB disconnection to SCSI RDMA Protocol (SRP) target was able to be reconnected but since SRP Target driver is designed to work directly on top of OpenFabrics OFED or Infiniband drivers in Linux kernel tree. I guess the fix could be a upgrading and patching of the Kernel or OFED drivers.

    Still noting this issue.