| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Until now, if two copies of one OVS daemon started up at the same time,
then due to races in pidfile creation it was possible for both of them to
start successfully, instead of just one. This was made worse when a
previous copy of the daemon had died abruptly, leaving a stale pidfile.
This commit implements a new pidfile creation and removal protocol that I
believe closes these races. Now, a pidfile is asserted with "link" instead
of "rename", which prevents the race on creation, and a stale pidfile may
only be deleted by a process after it has taken a lock on it.
This may solve mysterious problems seen occasionally on vswitch restart.
I'm still puzzled by these problems, however, because I don't see anything
in our tests cases that would actually cause two copies of a daemon to
start at the same time, which as far as I can see is a necessary
precondition for the problem.
|
| |
|
|
|
|
|
|
|
|
|
|
| |
Until now, it has been the responsibility of an individual daemon to call
die_if_already_running() at an appropriate time. A long time ago, this
had to happen *before* daemonizing, because once the process daemonized
itself there was no way to report failure to the process that originally
started the daemon. With the introduction of daemonize_start(), this is
now possible, but we haven't been taking advantage of it.
Therefore, this commit integrates the die_if_already_running() call into
daemonize_start() and deletes the calls to it from individual daemons.
|
| |
|
|
|
| |
It seems possible that a signal coming in at the wrong time could confuse
this code. It's always best to loop on EINTR.
|
| |
|
|
|
| |
If a daemon doesn't start, we need to know why. Being able to
consistently consult the log to find out is helpful.
|
| |
|
|
| |
This commit adds a few initial users but more are coming up.
|
| |
|
|
| |
This will acquire a new user in an upcoming commit.
|
| | |
|
| | |
|
| |
|
|
|
| |
PRINTF_FORMAT is more portable than raw __attribute__. We use it
elsewhere, I don't know why it wasn't used here.
|
| |
|
|
|
|
| |
This matches what xmalloc() does. It will be handled better by a monitor
process (created with --monitor), which will restart the child instead of
exiting.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
ovs-xapi-sync is supposed to always keep external-ids:iface-id up to date,
but in fact it would only set it when an interface initially appeared. If
the interface quickly disappeared and reappeared, then it failed to notice
that iface-id had changed or disappeared. This happens in practice on
Citrix XenServer, where VM "tap" devices often disappear and then reappear
almost immediately during VM boot. This commit fixes the problem.
This also fixes the similar problem for external-ids:bridge-id in Bridge
records. Bridges aren't ordinarily destroyed and re-created quickly, so
this problem might never have manifested in practice for bridges.
Many thanks to Reid Price <reid@nicira.com> for identifying the problem
and supplying an initial fix.
Bug #5239.
Reported-by: Henrik Amren <henrik@nicira.com>
|
| |
|
|
|
| |
Scanning the dsts twice seems may be a little more efficient than
partitioning it, and it now seems more straightforward to me.
|
| |
|
|
|
| |
In my opinion this is easier to understand than the way that these
two logically separate steps were previously entangled.
|
| |
|
|
|
|
|
| |
The following commit will need to iterate over a set of "struct
dst"s, obtaining the iface for each. It could look them up using
the hash table that indexes over dp_ifidx, but it's easier if we
simply store the iface pointer directly.
|
| |
|
|
| |
If it doesn't exist then it can't have the wrong value.
|
| |
|
|
|
|
| |
This removes over 1000 lines of code from bridge.c and will make it
easier to moving the bonding implementation into ofproto as part of
future development.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
The code that enables and disables bond slaves was a bit of a mess:
* Disabling a slave could recursively enable a different slave.
* Processing a flow could enable a slave.
This commit gets rid of both of those properties, which made it difficult
to reason about the code paths along which slaves would be enabled and
disabled.
Bug #5121.
|
| |
|
|
| |
It's quite clear that we don't support double tagging now.
|
| | |
|
| |
|
|
|
| |
I don't know why this was declared as 8 bytes long but I only see 6
actually in use, as one would expect of an Ethernet address.
|
| |
|
|
|
| |
ofproto_create() also calls dpif_flow_flush() very soon afterward. This
seems more clearly in ofproto's domain anyhow.
|
| |
|
|
| |
This makes it easier to pass configuration between modules.
|
| |
|
|
|
|
| |
There's no reason that I can see to maintain this information in struct
port and struct iface. It's redundant, since the lacp implementation
maintains the same information.
|
| | |
|
| |
|
|
|
|
|
|
|
| |
Only the first 6 bytes (ETH_ADDR_LEN) of the 'sys_id' argument are used,
but the prototype declared it as an array of 8 bytes. This has no effect
on the generated code--the declared size of an array parameter is
irrelevant--but it is misleading.
Also, add 'const' since the array is not modified.
|
| | |
|
| | |
|
| | |
|
| |
|
|
|
| |
This allows callers to add a VLAN header to the composed packet and send
it out on a VLAN without copying the whole payload.
|
| |
|
|
| |
This will soon be used in the upcoming bond library.
|
| |
|
|
|
|
|
|
|
| |
The second call to ofpbuf_put_zeros() could cause the 'eth' pointer to
be invalidated.
It appears that this does not fix a real bug because the existing callers
all preallocate 128 bytes of tailroom, but the interface doesn't document
that requirement.
|
| |
|
|
| |
The following commit will introduce the first use of snap_compose().
|
| |
|
|
|
| |
The new implementation of the bonding code expects to be able to send
packets on netdevs using netdev_send(). This implements it.
|
| |
|
|
|
|
| |
The new implementation of the bonding code expects to be able to
send packets using netdev_send(). This makes it possible for
Linux netdevs.
|
| |
|
|
|
|
|
|
|
| |
Sometimes lockfile will emit a message saying that it took a little while
to get the lock, which caused spurious test failures. This commit
suppresses the message. With this change, I was able to run these tests
continuously for some time without failures.
This was a bug in the testsuite, not in the code under test.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Before this (and the previous) patch, whenever cfm_configure was
called it would set the fault_timer to expired. Thus, the next
call to cfm_run would notice a lack of CCM reception and trigger a
faulted status. This is a bug in and of itself, but normally would
not be a big deal because cfm_configure should only be called
infrequently (when the database changes). However due to an
unrelated bug, cfm_configure() was getting called approximately once
per second. This resulted in all monitors showing faults all of
the time.
This patch fixes the problem by not expiring the timer at
cfm_configure(). Instead it gives it the appropriate
fault_interval amount of time to miss heartbeats.
Bug #5244.
|
| |
|
|
|
|
|
|
| |
Calling cfm_configure often could cause timers to be reset
resulting in unexpected behavior. This commit only updates when
cfm configuration actually changed.
Bug #5244.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
When ovsdb-server reads a database file that is corrupted at the
transaction level (that is, the transaction is valid JSON and has the
correct SHA-1 hash, but it does not describe a valid database transaction),
then ovsdb-server should truncate it and overwrite it by valid
transactions. However, until now, it didn't. Instead, it would keep the
invalid transaction and possibly every transaction in the database file
(depending on in what way the transaction was invalid), which would just
cause the same trouble again the next time the database was read.
This fixes the problem. An invalid transaction will be deleted from the
database file at the first write to the database.
Bug #5144.
Bug #5149.
|
| |
|
|
|
|
|
|
|
|
| |
When ovsdb-server reads a database that is corrupted at the log level
(that is, when ovsdb_log detects the corruption by checking the SHA-1 hash
of the record or JSON parser error reporting), then writing to the database
should discard the corrupted data and thereby fix the problem for future
ovsdb-server runs.
This already worked OK. This just adds an extra test.
|
| |
|
|
|
|
|
| |
If there's database corruption then it indicates that something went wrong,
e.g. the machine was powered-off by power failure. It's definitely
something that the admin should know about. This sounds like an error to
me, so use that log level.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
When a strong reference to a non-root table is ephemeral, the database log
can contain inconsistencies. In particular, if the column in question is
the only reference to a row, then the row will be created in one logged
transaction but the reference to it will not be logged (because it is
ephemeral). Thus, any later occurrence of the row later in the log (to
modify it, to delete it, or just to reference it) will yield a transaction
error and reading the database will abort at that point.
This commit fixes the problem by forcing any column with a strong reference
to a non-root table to be persistent.
The change to ovsdb_schema_from_json() looks bigger than it really is: it
just swaps the order of two operations on the schema and updates their
comments. Similarly for the update to ovs.db.DbSchema.__init__().
Bug #5144.
Reported-by: Sujatha Sumanth <ssumanth@nicira.com>
Bug #5149.
Reported-by: Ram Jothikumar <rjothikumar@nicira.com>
|
| |
|
|
|
|
|
|
|
|
| |
This function only worked properly inside OVSDB itself, because that is
the only place where the 'refTable' member of ovsdb_base_type is set.
Both inside and outside OVSDB, 'refTableName' is set for reference types,
so it's better to check for that.
This doesn't fix any existing bug because this function was only used
inside OVSDB until now.
|
| | |
|
| | |
|
| |
|
|
| |
Also deletes svec_split() since this was the only user.
|
| | |
|
| |
|
|
| |
Should be slightly cheaper than sorting a list (O(n) vs. O(n lg n)).
|
| | |
|
| | |
|
| |
|
|
|
|
| |
In each of the cases converted here, an shash was used simply to maintain
a set of strings, with the shash_nodes' 'data' values set to NULL. This
commit converts them to use sset instead.
|