summaryrefslogtreecommitdiff
path: root/datapath
Commit message (Collapse)AuthorAgeFilesLines
* datapath: IFF_BRIDGE_PORT is backported by Centos 5.6.Jesse Gross2011-09-211-1/+1
| | | | | | | | | | Some versions of Centos 5.6 backport the flag IFF_BRIDGE_PORT without the associated rx_handler changes, so this changes to use a version check since we really don't care about the actual symbol. Reported-by: Srinivasan Ramasubramanian <vrsrini@gmail.com> Signed-off-by: Jesse Gross <jesse@nicira.com>
* datapath: Cleanup actions.c:do_output().Jesse Gross2011-09-211-12/+11
| | | | | | | | The code for outputting a packet can be simplified a little and also modernized. There is no functional change. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Send to userspace errors shouldn't halt processing.Jesse Gross2011-09-211-1/+1
| | | | | | | | | | | | If we encounter an error when sending a packet to userspace due to an explicit action we stop processing further actions. This makes sense for things like push vlan, where to continue means outputting an incorrect packet. However, sending to userspace is more akin to outputting to a port, which does not halt further processing. For consistency, ignore errors in this case as well. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Correctly validate vport attributes on old kernels.Jesse Gross2011-09-201-2/+2
| | | | | | | | | | The vport policy for OVS_VPORT_ATTR_PORT_NO and OVS_VPORT_ATTR_TYPE are present only in the section for newer kernels. This means that on older kernels the length of these attributes are never checked anywhere but we go ahead and read from them anyways. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Remove check for shared skbs.Jesse Gross2011-09-201-2/+0
| | | | | | | | | We never allow shared skbs to be present inside of the OVS datapath but the presence of a check in the core makes this less clear. Since the check is very old and no longer relevant, drop it. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Fully initialize datapath before local port.Jesse Gross2011-09-202-25/+55
| | | | | | | | | | | | | | | It's possible to start receiving packets on a datapath as soon as the internal device is created. It's therefore important that the datapath be fully initialized before this, which it currently isn't. In particular, the fact that dp->stats_percpu is not yet set is potentially fatal. In addition, if allocation of the Netlink response failed it would leak the percpu memory. This fixes both problems. Found by code inspection, in practice the datapath is probably always done initializing before someone can send a packet on it. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Correctly set error code in queue_userspace_packets().Jesse Gross2011-09-191-1/+4
| | | | | | | | | | In a few places in queue_userspace_packets() when we encounter an error, we don't actually set the 'err' variable. Although we free the packets we don't correctly account for these packets as being lost. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Hardcode vport multicast group ID on older kernels.Ethan Jackson2011-09-161-1/+10
| | | | | | | | | | | Older kernels do not advertise the multicast groups of families when requested by userspace. As a workaround, this patch hardcodes the multicast group ID of the ovs_vport family on these kernels. Userspace will be able to fall back to this hardcoded value if the standard mechanism is unavailable. Signed-off-by: Ethan Jackson <ethan@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Always use generic stats for devices (vports)Pravin Shelar2011-09-1511-221/+71
| | | | | | | | | | | | | | | Currently ovs is using device stats for Linux devices and count them itself in other situations. This leads to overlap with hardware stats, inconsistencies, etc. It's much better to just always count the packets flowing through the switch and let userspace do any merging that it wants. Following patch removes vport->get_stats() interface. vport-stat is changed to use new `struct ovs_vport_stat` rather than rtnl_link_stats64. Definitions of rtnl_link_stats64 is removed from OVS. dipf_port->stat is also removed as aggregate stats are only available at netdev layer. Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* Set MTU in userspace rather than kernel.Justin Pettit2011-09-158-101/+0
| | | | | | | | | | | | Currently the kernel automatically sets the MTU of any internal interfaces to the minimum of all attached interfaces because the Linux bridge does this. Userspace can do this with more knowledge and flexibility. Feature #7323 Signed-off-by: Justin Pettit <jpettit@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Fix tunnel lookupPravin Shelar2011-09-141-1/+1
| | | | | | | | Attached patch fixes tunnel lookup to do correct port comparison. This bug is introduced by commit 3544358aa5960b148bc31435a0062e9392530ec2 Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Set vport in skb when executed from userspace.Jesse Gross2011-09-141-0/+5
| | | | | | | | | | Currently, the OVS_CB(skb)->vport member is never initialized for packets coming from userspace. This means that they can never be sampled by sFlow and generally violates our principle that userspace packets should be made to look the same as others. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Pravin Shelar <pshelar@nicira.com>
* datapath: Strip down vport interface : OVS_VPORT_ATTR_MTUPravin Shelar2011-09-121-8/+0
| | | | | | | | | | | | | | There is no need to have vport attribute MTU (OVS_VPORT_ATTR_MTU) as linux net-dev-ioctl can be used to get/set MTU for linux device. Following patch removes OVS_VPORT_ATTR_MTU from datapath protocol. This patch also adds netdev_set_mtu interface. So that MTU adjustments can be done from OVS userspace. get_mtu() interface is also changed, now get_mtu() returns EOPNOTSUPP rather than returning 0 and setting *pmtu to INT_MAX in case there is no MTU attribute for given device. Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: add key support to CAPWAP tunnelValient Gough2011-09-094-45/+265
| | | | | | | | | | | | | Add tunnel key support to CAPWAP vport. Uses the optional WSI field in a CAPWAP header to store a 64bit key. It can also be used without keys, in which case it is backward compatible with the old code. Documentation about the WSI field format is in CAPWAP.txt. Signed-off-by: Valient Gough <vgough@pobox.com> [horms@verge.net.au: Various minor fixes (v4.1)] Signed-off-by: Simon Horman <horms@verge.net.au> [jesse: Additional parsing fixes] Signed-off-by: Jesse Gross <jesse@nicira.com>
* datapath: Improve kernel hash tablePravin Shelar2011-09-0915-780/+390
| | | | | | | | | | | | Currently OVS uses its own hashing implmentation for hash tables which has some problems, e.g. error case on deletion code. Following patch replaces that with hlist based hash table which is consistent with other kernel hash tables. As Jesse suggested, flex-array is used for allocating hash buckets, So that we can have large hash-table without large contiguous kernel memory. Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: VLAN actions should use push/pop semanticsPravin Shelar2011-09-096-32/+188
| | | | | | | | | | | | | | Currently the kernel vlan actions mirror those used by OpenFlow 1.0. i.e. MODIFY and STRIP. More flexible approach is to have an action to push a tag and pop a tag off, so that it can handle multiple levels of vlan tags. Plus it aligns with newer version of OpenFlow. As this patch replaces MODIFY with PUSH semantic, action mapping done in userpace is fixed accordingly. GSO handling for multiple levels of vlan tags is also added as Jesse suggested before. Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Fix br_nlmsg_sizePravin Shelar2011-09-091-1/+0
| | | | | | | | | I missed this in last vport iflink patch. As IFLA_LINK is not be passed in netlink msg there is no need to allocate space for it. Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Allow a packet with no input port to omit OVS_KEY_ATTR_IN_PORT.Ben Pfaff2011-09-082-9/+9
| | | | | | | | | | | | | | | | | | When ovs-vswitchd executes actions on a synthesized packet, that is, on a packet that is not being forwarded from any particular port but is being generated by ovs-vswitchd itself or by an OpenFlow controller (using a OFPT_PACKET_OUT message with an in_port of OFPP_NONE), there is no good choice for the in_port to pass to the kernel in the flow in the OVS_PACKET_CMD_EXECUTE message. This commit allows ovs-vswitchd to omit the in_port entirely in this case. This fixes a bug in OFPT_PACKET_OUT: using an in_port of OFPP_NONE would cause the packet to be dropped by the kernel, since that's an invalid input port. Signed-off-by: Ben Pfaff <blp@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com> Reported-by: Aaron Rosen <arosen@clemson.edu>
* datapath: Calculate flow hash after extracting metadata.Jesse Gross2011-09-081-1/+2
| | | | | | | | | | | | | When we execute a packet from userspace we first extract the header fields from the packet and then add supplied metadata. However, we compute the hash of the packet in between these two steps despite the fact that the metadata can affect the hash. This can lead to two separate hashes for packets of the same flow. Found by code inspection, not an actual real-world problem. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* Strip down vport interface : iflinkPravin Shelar2011-09-086-48/+1
| | | | | | | | Remove iflink from vport interface. iflink is not used anywhere in OVS. So there is not need to have iflink as vport attribute. Signed-off-by: Pravin B Shelar <pshelar@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: genl_notify() on port disappearances.Ethan Jackson2011-09-013-4/+23
| | | | | | | | | | | Before this patch, if a vport detached itself from the datapath without interaction from userspace, rtnetlink notifications would be sent, but genl notifications would not. Feature #6809. Signed-off-by: Ethan Jackson <ethan@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Disable LRO from userspace instead of the kernel.Justin Pettit2011-08-282-31/+31
| | | | | | | | | | | | | | Whenever a port is added to the datapath, LRO is automatically disabled. In the future, we may want to enable LRO in some circumstances, so have userspace disable LRO through the ethtool ioctls. As part of this change, the MTU and LRO checks are moved to netdev-vport's send(), which is where they're actually needed. Feature #6810 Signed-off-by: Justin Pettit <jpettit@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Remove unneeded { } around single statement.Ben Pfaff2011-08-221-2/+1
| | | | | | | I noticed this looking around at other code. Signed-off-by: Ben Pfaff <blp@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Use "OVS_*" as opposed to "ODP_*" for user<->kernel interactions.Justin Pettit2011-08-1918-496/+496
| | | | | | | | | | | | | The prefix "ODP_*" is not overly descriptive in the context of the larger Linux tree. This commit changes the prefix to "OVS_*" for the userpace to kernel interactions. The userspace libraries still use "ODP_" in many of their interfaces since it is more descriptive in the OVS oeuvre. Feature #6904 Signed-off-by: Justin Pettit <jpettit@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Correct comment for vport_add().Justin Pettit2011-08-181-1/+1
| | | | | | | | The comment describing vport_add() incorrectly stated that the function added the vport to the datapath. It is the responsibility of the caller to do that. Signed-off-by: Justin Pettit <jpettit@nicira.com>
* ovs-dpctl: Show number of flowsSimon Horman2011-08-031-0/+3
| | | | | | | | | | | | | | | | | Expose the number of flows present in a datapath to user-space and to users via ovs-dpctl show. e.g.: ovs-dpctl show br3 system@br3: lookups: frags:0, hit:0, missed:0, lost:0 flows: 0 ... Signed-off-by: Simon Horman <horms@verge.net.au> [Jesse: Add same logic to userspace datapath.] Signed-off-by: Jesse Gross <jesse@nicira.com>
* datapath: Allow the number of hash entries to exceed TBL_MAX_BUCKETSSimon Horman2011-08-022-10/+14
| | | | | | | | | | | | | | | | * If the number of entries in a table exceeds the number of buckets that it has then an attempt will be made to resize the table. * There is a limit of TBL_MAX_BUCKETS placed on the number of buckets of a table. * If this limit is exceeded keep using the existing table. This allows a table to hold more than TBL_MAX_BUCKETS entries at the expense of increased hash collisions. Signed-off-by: Simon Horman <horms@verge.net.au> Signed-off-by: Jesse Gross <jesse@nicira.com>
* datapath: Allow table to expand to have TBL_MAX_BUCKETS bucketsSimon Horman2011-08-021-1/+1
| | | | | | | | This resolves what appears to be a logic error whereby the maximum number of buckets is limited to only half of TBL_MAX_BUCKETS. Signed-off-by: Simon Horman <horms@verge.net.au> Signed-off-by: Jesse Gross <jesse@nicira.com>
* datapath: Backport flex_arrays.Jesse Gross2011-07-285-0/+491
| | | | | | | | | flex_arrays didn't exist at all until 2.6.30, weren't exported to modules until 2.6.38, and performed poorly until 3.0, so this backports the functionality to older kernels. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Don't pass __GFP_ZERO to kmalloc on older kernels.Jesse Gross2011-07-281-0/+17
| | | | | | | | | | | | | On new kernels kzalloc() is simply a wrapper around kmalloc with the addition of the __GFP_ZERO flag. flex_arrays take advantage of this by expecting the user to just pass in this flag if they want the memory to be zeroed. However, before 2.6.23, kzalloc() was a function in its own right and kmalloc really didn't like receiving __GFP_ZERO. This overrides kmalloc() to intercept the flags and direct the call to the right function. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Backport reciprocal division.Jesse Gross2011-07-284-0/+52
| | | | | | | | The reciprocal division library did not exist until 2.6.20 and is not currently exported in any version, so this backports it. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* Datapath action should not refer to controllerpravin shelar2011-07-282-8/+8
| | | | | | | | ODP_ACTION_ATTR_CONTROLLER in the kernel actually sends packets to userspace, not the controller. To make it generic rename this action to ODP_ACTION_ATTR_USERSPACE. Signed-off-by: Pravin B Shelar <pshelar@nicira.com>
* datapath: Add missing targets to avoid failure on e.g. "make TAGS".Ben Pfaff2011-07-261-3/+22
| | | | | | | | | | | | Automake invokes a number of targets recursively, including in datapath/linux, so we need to define all those targets or get an error from "make" when those targets are invoked, e.g. when "make TAGS" is run. This commit adds a no-op target to the main datapath/linux Makefile for each recursive target listed in the "Third-Party Makefiles" section of the Automake manual, in the order listed there, fixing the problem. Signed-off-by: Ben Pfaff <blp@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: An expanded table should be larger than its predecessorSimon Horman2011-07-121-1/+1
| | | | | | | | | | | | | | | | | | This resolves what appears to be a think-o in tbl_expand() * Old Logic: Always create tables with TBL_MIN_BUCKETS buckets * New Logic: Create tables twice as big as their predecessor When sending 10,000 flows through ovs-vswitchd: * Old Logic: CPU bound in tbl_lookup(), significant packet loss * New Logic: ~10% of one core used, negligible packet loss Tested with an Intel E5520 @ 2.27GHz, flows from an ethernet device to to a dummy interface with no address configured. Signed-off-by: Simon Horman <horms@verge.net.au> Signed-off-by: Jesse Gross <jesse@nicira.com>
* tunneling: Force selection of an IP ID with GRE.Jesse Gross2011-06-301-1/+4
| | | | | | | | | | | | | | | | | | By default we set the DF bit on tunneled packets because we want to get path MTU discovery from the underlying network. In turn this causes Linux to leave the IP ID as 0 because it believes that fragmentation can never occur. However, with GRE fragmentation is still possible because we may get a large packet to be encapsulated and let the local IP stack do fragmentation. As long as packets are kept in order fragments are not misassociated and everything works fine. However, if there is reordering in the underlying network then packets can become corrupted. This forces selection of an IP ID for GRE packets to avoid misassociation. Bug #6128 Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Support Linux 3.0Simon Horman2011-06-291-2/+2
| | | | | | | This trivially supports linux 3.0 by incrementing the version check. Signed-off-by: Simon Horman <horms@verge.net.au> Signed-off-by: Jesse Gross <jesse@nicira.com>
* datapath: Rename linux-2.6 and compat-2.6 directories.Jesse Gross2011-06-2464-66/+65
| | | | | | | | The linux-2.6 and compat-2.6 directories apply equally to the upcoming Linux 3.0 release, so this drops the 2.6 suffix and updates Makefiles. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Add missing header.Jesse Gross2011-06-231-0/+1
| | | | | | | | | The internal dev vport really needs hardirq.h but doesn't depend directly on it and has relied on it being included from other sources. Recent kernels broke this, so explicitly add the header. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* configure: Do not reject Linux 3.0 at configure time.Ben Pfaff2011-06-221-4/+0
| | | | | Until now, the configure script has rejected any version of Linux other than 2.6. In preparation for Linux 3.0, this allows newer versions also.
* configure: Remove "26" from Linux variable names.Ben Pfaff2011-06-222-2/+2
| | | | | | | | | OVS used to support Linux 2.4 and Linux 2.6, but now it only supports Linux 2.6. Linux 3.0 is coming up, and it's just an evolution of 2.6, so OVS should stop referring to it as "2.6". This takes a first step by removing "26" from internal variable names. There should be no user-visible changes.
* datapath: Use consume_skb() on non-errors.Jesse Gross2011-06-164-12/+16
| | | | | | | | | | | It's possible to trace kfree_skb() call sites to find out where packets are getting dropped. Situations where kfree_skb() does not actually indicate an error adds additional noise, so use consume_skb() instead to avoid tracing non-errors. Suggested-by: Ben Pfaff <blp@nicira.com> Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Backport consume_skb().Jesse Gross2011-06-161-0/+4
| | | | | | | | Kernels before 2.6.30 did not implement consume_skb() although RHEL backports it. For other kernels, this provides a backport. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Further mirror checksum offloading state on old kernels.Jesse Gross2011-06-167-133/+257
| | | | | | | | | | | | | | | | | | Older kernels (those before 2.6.22) rely on implicit assumptions to determine checksum offloading status. These assumptions tend to break down when doing switching because it sits in the middle of the transmit and receive path. Newer kernels deal with this problem by adding more explicit information about how to checksum. This replicates that behavior by mirroring the state from newer kernels in private OVS storage on the kernels that lack it. On ingress and egress we then map that state onto the appropriate location for the given kernel and can consistently manipulate it within OVS. Some of this was already done for the checksum type but this makes it more robust and expands it to the checksum start and offset as well. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Drop set_skb_csum_bits().Jesse Gross2011-06-161-18/+0
| | | | | | | | | | | Various older kernels have had different bugs with copying checksum state when a complete copy of a packet is made. However, it is not actually necessary to make these copies and all occurrences have now been removed. Therefore, we can also remove the workarounds to deal with these bugs. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* tunneling: Avoid extra copying if expanding headroom.Jesse Gross2011-06-161-25/+8
| | | | | | | | | | | | Currently if we need additional headroom before encapsulating a packet a clone is made before expanding headroom or if we are just trying to make the headroom writable then we copy both the struct sk_buff and the paged data. Both of these are unnecessary and we end up freeing the original copy. We can remove these copies and simplify the code by just expanding the linear data area. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Simplify make_writable().Jesse Gross2011-06-161-81/+79
| | | | | | | | | | | | | | The current implementation of make_writable() is both overly complex and unnecessarily aggressive about copying data. We can improve performance by only making a copy of the data if someone else holds a reference to the portion of the data that we want to modify. This means that if a clone is held by the TCP stack for retransmission then we do not need to make a copy if we are changing the IP header because it will get regenerated on retransmit anyways. Even when it is necessary to copy we avoid duplicating struct sk_buff. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Use strip_vlan() for modify_vlan_tci().Jesse Gross2011-06-161-22/+8
| | | | | | | | | | | | The sematics for setting a vlan tag are to modify the existing tag if one exists. This can be expressed as removing the existing tag first and then adding a new one. This simplifies the code by not requiring two copies of the logic that manipulates non-accelerated vlans and should not make a performance difference because the vlan tag is contained in a single cache line. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>
* datapath: Check for supported kernel versions.Jesse Gross2011-06-131-0/+5
| | | | | | | | | Most of the time kernels older or newer than the ones we support simply fail to compile. However, sometimes they appear to succeed but then cause problems later on. This explicitly checks for supported versions at compile time. Signed-off-by: Jesse Gross <jesse@nicira.com>
* Remove NXAST_DROP_SPOOFED_ARP action.Justin Pettit2011-06-092-36/+1
| | | | | | | | | The NXAST_DROP_SPOOFED_ARP action has been deprecated in favor of defining flows using the NXM_NX_ARP_SHA flow match for a while. This commit removes it. Signed-off-by: Justin Pettit <jpettit@nicira.com> Acked-by: Jesse Gross <jesse@nicira.com>
* datapath: Remove redundant nw_ prefix from fields in flow key.Jesse Gross2011-06-083-46/+46
| | | | | | | | | | | The fields of the kernel flow key are now grouped by protocol rather than using generic names. The containing structures describe the category, so it is no longer necessary to use prefixes. Most of these prefixes have been removed but nw_proto and nw_tos have retained them. This renames the fields for consistency and brevity. Signed-off-by: Jesse Gross <jesse@nicira.com> Acked-by: Ben Pfaff <blp@nicira.com>