summaryrefslogtreecommitdiff
path: root/lib/netdev-linux.c
Commit message (Collapse)AuthorAgeFilesLines
...
* netdev-linux: Add capability to get stats from vport layer.Jesse Gross2010-06-101-23/+36
| | | | | | | The vport layer has the ability to track stats using 64-bit counters, even if the kernel is only 32-bit. This first attempts to collect stats from these counters if they are available and otherwise falls back to the normal Linux interfaces.
* netdev-linux: Give tap FD to first opener.Jesse Gross2010-06-011-1/+9
| | | | | | | | | | Tap devices can have two FDs that allow transmit and receive from different perspectives. We previously would always share one of the FDs among all openers. However, this is confusing to some users (primarily the DHCP client) which expect tap devices to behave like any other device. Now we give the tap FD to the first opener, which knows that it has opened a tap device, and a normal system FD to everyone else for consistency.
* netdev-linux: Fix tap device stats.Jesse Gross2010-06-011-19/+15
| | | | | | | | For tap and internal devices we swap the transmit and receive stats to appear consistent with other devices. However, the check whether to store the stats in a temporary location before the swap did not include tap devices, which lead to the use of uninitialized memory when the swap occured.
* netdev-linux: Quiet down ingress policing.Jesse Gross2010-05-191-1/+1
| | | | | | | If we attempt to remove ingress policing and receive "invalid argument" it means that policing isn't compiled into the kernel. If it isn't compiled in then accept that policing has been successfully removed.
* patch: Remove veth driver.Jesse Gross2010-05-181-208/+0
| | | | | | Now that we have a new patch implementation, remove the veth driver and its userspace components. Then rename 'patchnew' to 'patch'. The new implementation is a drop-in replacement for the old one.
* netdev-linux: Optimize removing policing from an interface.Ben Pfaff2010-05-051-14/+42
| | | | | | | | | | It is very expensive to start a subprocess and, especially, to wait for it to complete. This replaces the most common subprocess operation in netdev_linux_set_policing() by a Netlink socket operation, which is much faster. Without this and the other netdev-linux commits, my 1000-interface test case runs in 1 min 48 s. With them, it runs in 25 seconds.
* netdev-linux: Cache policing values.Ben Pfaff2010-05-051-6/+26
| | | | | Without this and the following netdev-linux commits, my 1000-interface test case runs in 1 min 48 s. With them, it runs in 25 seconds.
* netdev-linux: Factor out removing policing.Ben Pfaff2010-05-051-12/+20
| | | | This is duplicated code that the following commit will rewrite.
* netdev-linux: Factor out obtaining an RTNL socket.Ben Pfaff2010-05-051-9/+28
| | | | | Another function needs this same functionality in an upcoming commit, so factor this into a new function get_rtnl_sock().
* Update fake bond devices' statistics with the sum of bond slaves' stats.Ben Pfaff2010-04-191-23/+76
| | | | | | | | Needed by XAPI to accurately report bond statistics. Ugh. Bug NIC-63.
* tunneling: Remove old GRE implementation.Jesse Gross2010-04-191-531/+8
| | | | | | The new GRE implementation provides a complete drop in replacement for the old Linux based implementation. Therefore, remove the old implementation and rename "grenew" to "gre".
* netdev-linux: Don't free a member of a struct.Jesse Gross2010-04-191-1/+1
| | | | | | | We allocate struct netdev_linux which contains struct netdev but free the netdev. In practice this makes no difference because the netdev is the first member of the struct but we should be correct anyways.
* netdev-linux: Check notifications are for netdev-linux device.Jesse Gross2010-04-191-8/+21
| | | | | | When receiving a change notification from rtnetlink we checked whether a netdev of that name existed and if so tried to handle it. This also checks that the type of the device is one handled by netdev-linux.
* netdev: Add support for "patch" typeJustin Pettit2010-04-151-2/+193
| | | | | | | | | | | | | | | | | | | This commit introduces a new netdev type called "patch". A patch is a pair of interfaces, in which frames sent through one of the devices pop out of the other. This is useful for linking together datapaths. A patch's only argument on creation is "peer", which specifies the other side of the patch. A patch must be created in pairs, so a second netdev must be created with the "name" and "peer" values reversed. The current implementation is built using veth devices. Further, it's limited to the veth devices which support configuration through sysfs. This limits the ability to use a "patch" on 2.6.18 kernels using the veth device we include (read: flavors of XenServer 5.5). In the not too distant future, the implementation will be modified to use the new kernel port abstraction introduced by Jesse Gross's forthcoming GRE work. At that point, patch devices will work on any Linux platform supported by OVS.
* gre: Add support for path MTU discovery.Jesse Gross2010-03-051-2/+11
| | | | | | | | | | | | | | | | | | | | | | | This allows path MTU discovery to properly work when used with bridging. While there was previously support for PMTUD it used the kernel's IP stack. This works fine for routing but when bridging it is possible that a complete network is operating over the bridge that the kernel has no knowledge of and the ICMP fragmentation needed packets are lost. When a packet arrives that is above the MTU of the tunnel, an ICMP message is synthesized and send back on the device that the original packet came from. This does not rely on the kernel IP stack and is therefore independent of the routing table. Both IPv4 and IPv6 are supported, including over VLANs. Other types of packets that are over the MTU are encapsulated and the outer packets are fragmented. This entire functionality is a layer violation since bridging operates at layer 2 and fragmentation is a function of layer 3. For this reason it is possible to disable PMTUD, which will provide complete transparency but will cause the outer IP packets to be fragmented.
* gre: Allow ToS on outer packet to be configured.Jesse Gross2010-03-051-1/+5
| | | | | | When creating a GRE tunnel, it is now possible to either set the ToS of the outer packet to a fixed value or copy it from the inner packet.
* gre: Always set TTL on outer packet to 64.Jesse Gross2010-03-051-1/+2
| | | | | | | | | | | | | | | | Currently the TTL is copied from the inner packet of the tunnel to the outer packet if the inner packet is IP. This is good if your GRE packets might make it into the input of your device but bad if you want to be fully transparent. This also resolves an inconsistency between tunnels set up using the ioctl and using Netlink. The ioctl version would force PMTUD on if a fixed TTL is set as a backup way to prevent loops but it never made it over to the newer Netlink code so obviously no one cares too much about it. This removes it to provide consistency and transparency. Basically, don't create loops and you will be happy.
* Merge "master" into "next".Ben Pfaff2010-02-111-10/+10
|\ | | | | | | | | The main change here is the need to update all of the uses of UNUSED in the next branch to OVS_UNUSED as it is now spelled on "master".
| * Rename UNUSED macro to OVS_UNUSED to avoid naming conflict.Ben Pfaff2010-02-111-4/+4
| | | | | | | | Requested by Jean Tourrilhes <jt@hpl.hp.com>.
* | netdev-linux: Avoid fiddling with indeterminate data.Ben Pfaff2010-02-111-4/+2
| | | | | | | | | | | | | | | | | | | | If we are using netlink to get stats and get_ifindex() fails, then for an internal network device we will then swap around a bunch of indeterminate (uninitialized) data values. That won't hurt anything--the caller will still set them to all-1-bits due to the error--but it still seems wrong. So this commit avoid it. Found using Clang (http://clang-analyzer.llvm.org/).
* | netdev-linux: Use the netdev list of devices instead of cachemap.Jesse Gross2010-01-181-9/+15
| | | | | | | | | | | | We previously maintained a list of open devices inside of the linux netdev. Since the netdev library now maintains this list, it is better to use that list instead of our own.
* | netdev-linux: Avoid potential issues with unset FD.Jesse Gross2010-01-181-3/+2
| | | | | | | | | | | | Never close the file descriptor if it is 0, since it is never a valid FD in this context. Also initialize the FD to -1 so that it is never set to a valid but incorrect value.
* | netdev-linux: Properly store netdev_dev pointer for RTNL callbacks.Jesse Gross2010-01-161-1/+1
| | | | | | | | | | | | We were storing a struct netdev_dev_linux ** instead of a netdev_dev_linux * in the cache map. This prevented the cache from being invalidated on changes such as link status.
* | netdev: Increase default ingress policing burst sizeJustin Pettit2010-01-151-2/+2
| | | | | | | | | | | | | | The default burst rate was 10Kb. This increases it to 1000kb, since we were having problems getting traffic through at 10kb. A better value probably exists between these two points, but that will require additional experimentation.
* | netdev-linux: Don't close(0) when closing an ordinary netdev.Ben Pfaff2010-01-151-1/+2
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Calling close(0) at random points is bad. It means that the next call to socket() or open() returns fd 0. Then the next time a netdev gets closed, that socket or file fd gets closed too, and you end up with weird "Bad file descriptor" errors. Found by installing the following as lib/unistd.h in the source tree: #ifndef UNISTD_H #define UNISTD_H 1 #include <stdlib.h> #include_next <unistd.h> #undef close #define close(fd) rpl_close(fd) static inline int rpl_close(int fd) { if (!fd) { abort(); } return (close)(fd); } #endif
* | netdev-linux: Cleanup tap netdev.Jesse Gross2010-01-151-69/+35
| | | | | | | | | | | | | | | | | | | | | | | | TAP devices need to be treated slightly differently from other other devices because they cannot be opened multiple times. Instead we open them once and share the file descriptor. This means that if the netdev is opened multiple times one reader can drain the buffers of another. While this is a deviation from the normal convention, it does not impact current or planned users. In addition, this cleans up some confusion between the file descriptor for tap devices versus other FD's.
* | gre: Add support for destroying GRE devices.Jesse Gross2010-01-151-24/+138
| | | | | | | | | | This allows GRE tunnel devices to be torn down on graceful exit of vswitch and cleaned up on restart for non-graceful exits.
* | netdev: Fully handle netdev lifecycle through refcounting.Jesse Gross2010-01-151-228/+232
| | | | | | | | | | | | | | | | | | | | This builds on earlier work that implemented netdev object refcounting. However, rather than requiring explicit create and destroy calls, these operations are now performed automatically based on the referenece count. This is important because in certain situations it is not possible to know whether a netdev has already been created. A workaround existed (which looked fairly similar to this paradigm) but introduced it's own issues. This simplifies and unifies the API.
* | netdev-linux: Fix aliasing error.Ben Pfaff2009-12-141-5/+5
| | | | | | | | | | | | The latest version of GCC flags a common socket convention as breaking strict-aliasing rules. This commit removes the aliasing and gets rid of the scary warning.
* | gre: Add userspace GRE support.Jesse Gross2009-12-071-21/+469
| | | | | | | | | | | | | | | | | | This implements the userspace portion of GRE on Linux. It communicates with the kernel module to setup tunnels using either Netlink or ioctls as appropriate based on the kernel version. Significant portions of this commit were actually written by Justin Pettit.
* | Merge "master" branch into "db".Ben Pfaff2009-12-021-15/+116
|\ \ | |/
| * netdev: Allow explicit creation of netdev objectsJustin Pettit2009-12-011-13/+102
| | | | | | | | | | | | | | | | | | | | This change adds netdev_create() and netdev_destroy() functions to allow the creation of network devices through the netdev library. Previously, network devices had to already exist or be created on demand through netdev_open(). This caused problems such as not being able to specify TAP devices as ports in ovs-vswitchd, which this patch fixes. This also lays the groundwork for adding GRE and VDE support.
| * netdev: New function netdev_get_ifindex().Ben Pfaff2009-11-231-0/+13
| | | | | | | | | | sFlow needs the ifindex of an interface, so this commit adds a function to retrieve it.
| * netdev: Really set output values to 0 on failure in netdev_get_features().Ben Pfaff2009-11-191-2/+1
| | | | | | | | | | | | | | | | The comment on netdev_get_features() claimed that all of the passed-in values were set to 0 on failure, but the implementation didn't live up to the promise. CC: Paul Ingram <paul@nicira.com>
* | Add new function xzalloc(n) as a shorthand for xcalloc(1, n).Ben Pfaff2009-11-041-1/+1
|/
* netdev-linux: Improve netdev_linux_set_etheraddr().Ben Pfaff2009-10-021-3/+11
| | | | | | | | | Fixes a bug whereby netdev_linux_set_etheraddr() would update the cached Ethernet address but not mark it valid. (This potentially wasted a system call later but wasn't harmful.) As an added optimization, don't set the Ethernet address at all if the new address is the same as the current address.
* netdev-linux: Return correct error codes on receive.Jesse Gross2009-10-021-2/+2
| | | | | | | netdev_linux_receive was returning positive error codes while the interface specifies that it should be returning negative errors. This difference causes a huge increase in (non-existant) packet processing with the userspace datapath.
* netdev-linux: Fix tap device using wrong FD.Jesse Gross2009-09-301-3/+5
| | | | | Tap devices were doing ioctls on the AF_INET socket, instead of the FD opened on the tap device.
* Merge citrix branch into master.Ben Pfaff2009-09-221-0/+3
|
* netdev-linux: Set missing cache validity bit.Jesse Gross2009-09-161-0/+2
| | | | | | | | | Whether a port is internal is cached to avoid requerying the kernel every time stats are requested. However, the cache vality bit was never being set so the cache wasn't used. This corrects that oversight. Thanks to Ben Pfaff for noticing.
* netdev: Swap transmit and receive stats on internal ports.Jesse Gross2009-09-141-5/+63
| | | | | | | Internal ports appear to have their transmit and receive stats swapped because from the kernel's point of view these ports are acting like the machine connected to the switch, not the switch itself. This swaps the stats for consistency with other ports.
* Merge citrix branch into master.Ben Pfaff2009-09-021-21/+101
|
* rtnetlink: Move into separate source and header file.Ben Pfaff2009-07-301-157/+12
| | | | | Now that rtnetlink isn't named similarly to netdev_linux, it might as well have its own source and header files to avoid confusing everyone.
* rtnetlink: Document.Ben Pfaff2009-07-301-0/+15
|
* netdev-linux: Rename "linux_netdev_*" to "rtnetlink_*".Ben Pfaff2009-07-301-31/+31
| | | | | | | | It was getting to be too confusing to have both netdev_linux_* functions and linux_netdev_* functions. Rename the latter to make the distinction more obvious. "rtnetlink" seems to be a fairly good name because that's what the kernel calls it, so the name will be familiar at least to people who know about rtnetlink.
* netdev: Implement an abstract interface to network devices.Ben Pfaff2009-07-301-40/+1580
| | | | | | | | | This new abstraction layer allows multiple implementations of network devices in a single running process. This will be useful, for example, to support network devices that are simulated entirely in the running process or that communicate with other processes over Unix domain sockets, etc. The reimplemented tap device support in this commit has not been tested.
* Introduce general-purpose ways to wait for dpif and netdev changes.Ben Pfaff2009-07-061-0/+183
The dpif and netdev code has had various ways to check for changes to dpifs and netdevs over the course of Open vSwitch development. All of these have been thus far fairly specific to the Linux implementation. This commit is the start of a more general API for watching for such changes. The dpif-related parts seem fairly mature and so they are documented, the netdev parts will probably need to change somewhat and so they are not documented yet.