diff options
| author | Ben Pfaff <blp@nicira.com> | 2011-01-24 14:59:57 -0800 |
|---|---|---|
| committer | Ben Pfaff <blp@nicira.com> | 2011-01-27 21:08:36 -0800 |
| commit | 856081f683d3e7d5b5fa07af4233d285eb205c47 (patch) | |
| tree | f626bca05553ff757f1810f517b6d928ec0b1c28 /lib/dpif.h | |
| parent | 36956a7d33c9ee204fcb184484a5aaacbd9ecef8 (diff) | |
| download | openvswitch-856081f683d3e7d5b5fa07af4233d285eb205c47.tar.gz | |
datapath: Report kernel's flow key when passing packets up to userspace.
One of the goals for Open vSwitch is to decouple kernel and userspace
software, so that either one can be upgraded or rolled back independent of
the other. To do this in full generality, it must be possible to change
the kernel's idea of the flow key separately from the userspace version.
This commit takes one step in that direction by making the kernel report
its idea of the flow that a packet belongs to whenever it passes a packet
up to userspace. This means that userspace can intelligently figure out
what to do:
- If userspace's notion of the flow for the packet matches the kernel's,
then nothing special is necessary.
- If the kernel has a more specific notion for the flow than userspace,
for example if the kernel decoded IPv6 headers but userspace stopped
at the Ethernet type (because it does not understand IPv6), then again
nothing special is necessary: userspace can still set up the flow in
the usual way.
- If userspace has a more specific notion for the flow than the kernel,
for example if userspace decoded an IPv6 header but the kernel
stopped at the Ethernet type, then userspace can forward the packet
manually, without setting up a flow in the kernel. (This case is
bad from a performance point of view, but at least it is correct.)
This commit does not actually make userspace flexible enough to handle
changes in the kernel flow key structure, although userspace does now
have enough information to do that intelligently. This will have to wait
for later commits.
This commit is bigger than it would otherwise be because it is rolled
together with changing "struct odp_msg" to a sequence of Netlink
attributes. The alternative, to do each of those changes in a separate
patch, seemed like overkill because it meant that either we would have to
introduce and then kill off Netlink attributes for in_port and tun_id, if
Netlink conversion went first, or shove yet another variable-length header
into the stuff already after odp_msg, if adding the flow key to odp_msg
went first.
This commit will slow down performance of checksumming packets sent up to
userspace. I'm not entirely pleased with how I did it. I considered a
couple of alternatives, but none of them seemed that much better.
Suggestions welcome. Not changing anything wasn't an option,
unfortunately. At any rate some slowdown will become unavoidable when OVS
actually starts using Netlink instead of just Netlink framing.
(Actually, I thought of one option where we could avoid that: make
userspace do the checksum instead, by passing csum_start and csum_offset as
part of what goes to userspace. But that's not perfect either.)
Signed-off-by: Ben Pfaff <blp@nicira.com>
Acked-by: Jesse Gross <jesse@nicira.com>
Diffstat (limited to 'lib/dpif.h')
| -rw-r--r-- | lib/dpif.h | 32 |
1 files changed, 24 insertions, 8 deletions
diff --git a/lib/dpif.h b/lib/dpif.h index 0a41b77b9..3e72539c1 100644 --- a/lib/dpif.h +++ b/lib/dpif.h @@ -92,19 +92,35 @@ int dpif_flow_dump_done(struct dpif_flow_dump *); int dpif_execute(struct dpif *, const struct nlattr *actions, size_t actions_len, const struct ofpbuf *); -/* Minimum number of bytes of headroom for a packet returned by dpif_recv() - * member function. This headroom allows "struct odp_msg" to be replaced by - * "struct ofp_packet_in" without copying the buffer. */ -#define DPIF_RECV_MSG_PADDING \ - ROUND_UP(sizeof(struct ofp_packet_in) - sizeof(struct odp_msg), 8) -BUILD_ASSERT_DECL(sizeof(struct ofp_packet_in) > sizeof(struct odp_msg)); -BUILD_ASSERT_DECL(DPIF_RECV_MSG_PADDING % 8 == 0); +/* A packet passed up from the datapath to userspace. + * + * If 'key' or 'actions' is nonnull, then it points into data owned by + * 'packet', so their memory cannot be freed separately. (This is hardly a + * great way to do things but it works out OK for the dpif providers and + * clients that exist so far.) + */ +struct dpif_upcall { + uint32_t type; /* One of _ODPL_*_NR. */ + + /* All types. */ + struct ofpbuf *packet; /* Packet data. */ + struct nlattr *key; /* Flow key. */ + size_t key_len; /* Length of 'key' in bytes. */ + + /* _ODPL_ACTION_NR only. */ + uint64_t userdata; /* Argument to ODPAT_CONTROLLER. */ + + /* _ODPL_SFLOW_NR only. */ + uint32_t sample_pool; /* # of sampling candidate packets so far. */ + struct nlattr *actions; /* Associated flow actions. */ + size_t actions_len; +}; int dpif_recv_get_mask(const struct dpif *, int *listen_mask); int dpif_recv_set_mask(struct dpif *, int listen_mask); int dpif_get_sflow_probability(const struct dpif *, uint32_t *probability); int dpif_set_sflow_probability(struct dpif *, uint32_t probability); -int dpif_recv(struct dpif *, struct ofpbuf **); +int dpif_recv(struct dpif *, struct dpif_upcall *); int dpif_recv_purge(struct dpif *); void dpif_recv_wait(struct dpif *); |
