summaryrefslogtreecommitdiff
path: root/src/rabbit_node_monitor.erl
Commit message (Collapse)AuthorAgeFilesLines
* Log nodedown_reason.bug25922Simon MacMullen2013-12-101-2/+4
|
* Fix stupiditybug25780Simon MacMullen2013-09-251-1/+2
|
* Don't go through the legacy code path when not dealing with a legacy file; ↵Simon MacMullen2013-09-231-10/+8
| | | | don't arbitrarily rewrite the disc nodes field of this file at startup.
* Silence the badarg on two near-simultaneous pauses.bug25710Simon MacMullen2013-08-131-4/+6
|
* await_cluster_recovery/0 should return ok.bug25700Simon MacMullen2013-08-071-1/+2
|
* Refresh branch from stableEmile Joubert2013-07-311-3/+8
|\
| * More sensible API for partitions, do not return errors.bug25651Simon MacMullen2013-07-041-3/+8
| |
* | s/VMware/GoPivotal/gSimon MacMullen2013-07-011-2/+2
|/
* Look for the rabbit process, not the rabbit application.Simon MacMullen2013-05-241-1/+1
|
* Switch pause_minority mode to making its decisions entirely based on ↵bug25471Simon MacMullen2013-04-221-2/+18
| | | | node-upness, not rabbit-application-upness and explain why.
* Move those functions to their own place, and replace the autoheal ↵Simon MacMullen2013-04-221-20/+29
| | | | all_nodes_up check with all_rabbit_nodes_up since it will depend on the rabbit application running to DTRT.
* I suppose it would be polite to add specs here.Simon MacMullen2013-04-181-0/+3
|
* First pass at splitting all the autoheal stuff out into a separate module.Simon MacMullen2013-04-171-188/+15
|
* Remove io:formatsSimon MacMullen2013-04-081-2/+0
|
* Deal with partial partitions.Simon MacMullen2013-04-081-13/+19
|
* Umm, order by *best* partition first!Simon MacMullen2013-04-081-2/+3
|
* WIP commit: start to deal with non-consistent partitions. all_partitions/1 ↵Simon MacMullen2013-03-281-16/+54
| | | | needs attention though.
* Merge defaultSimon MacMullen2013-03-271-9/+161
|\
| * CosmeticSimon MacMullen2013-03-271-1/+2
| |
| * Merge in default, add a touch more logging, and don't restart autoheal if ↵Simon MacMullen2013-03-271-19/+57
| |\ | | | | | | | | | it's already started.
| * | Simplistic way of selecting a winner: order by the number of connections, ↵Simon MacMullen2013-03-261-19/+28
| | | | | | | | | | | | then the size of the partition.
| * | Figure out a global view of partitions.Simon MacMullen2013-03-261-1/+21
| | |
| * | Add some loggingSimon MacMullen2013-03-261-3/+11
| | |
| * | First pass at autohealing.Simon MacMullen2013-03-191-11/+121
| | |
* | | Merged bug25499 into defaultEmile Joubert2013-03-271-1/+4
|\ \ \ | |_|/ |/| |
| * | Check if the rabbit process is running and thus avoid deadlocks in the ↵bug25499Simon MacMullen2013-03-221-1/+1
| | | | | | | | | | | | application controller.
| * | A node counts as down if it is not running rabbit in this case.Simon MacMullen2013-03-191-1/+4
| |/
* | Don't call rabbit_mnesia:cluster_nodes(all) as much, don't use floating ↵bug25491Simon MacMullen2013-03-251-9/+14
| | | | | | | | point if we don't have to.
* | Use a macroSimon MacMullen2013-03-251-1/+1
| |
* | Do the pinging in a separate process so we don't block the node monitor.Simon MacMullen2013-03-211-6/+18
| |
* | First attempt at pinging that doesn't really work well.Simon MacMullen2013-03-201-5/+22
|/
* Code comment about clearing partitionsbug25474Emile Joubert2013-03-181-2/+2
|
* If we have been partitioned, and we are now in the only remaining partition, ↵Simon MacMullen2013-03-121-3/+21
| | | | we no longer care about partitions - forget them. Note that we do not attempt to deal with individual (other) partitions going away, it's only safe to forget *any* of them when we have seen the back of *all* of them.
* Treat {inconsistent_database, running_partitioned_network, Node} as being ↵bug25486Simon MacMullen2013-03-121-2/+11
| | | | sort of like {node_up, Node, NodeType}. It's not perfect, but it's the best we're going to get.
* Oops. This was part of an (early, wrong) attempt at bug 25474 which got ↵Simon MacMullen2013-03-121-4/+3
| | | | committed as part of f1317bb80df9 (bug 25358) by mistake. Remove.
* Register the process name to make sure we only have one running at a time.bug25358Simon MacMullen2013-03-061-0/+3
|
* We don't need this - net_adm:ping/1 will not return pong for nodes that have ↵Simon MacMullen2013-03-051-4/+4
| | | | gone down before net_ticktime expires (it will hang for until then instead).
* Merge in default (umm, again)Simon MacMullen2013-03-051-5/+17
|\
| * Merged stable into defaultEmile Joubert2013-03-051-1/+1
| |\
| * | Allow subscribing to node down events from the node monitor.bug25472Simon MacMullen2013-03-011-5/+17
| | |
* | | Rename this thing, to make space for bug 25471Simon MacMullen2013-03-011-6/+12
| | |
* | | We no longer need two different death detectors since we no longer look at ↵Simon MacMullen2013-02-271-10/+1
| | | | | | | | | | | | Mnesia for majorityness.
* | | SimplifySimon MacMullen2013-02-271-3/+1
| | |
* | | Base the whole thing off net_adm:ping/1 - because we might see other nodes ↵Simon MacMullen2013-02-271-5/+10
| | | | | | | | | | | | come back but also be waiting (in the no-majority case, and RAM nodes). Better to detect they exist and come back than to stay stuck because they don't happen to be running Mnesia.
* | | When we lose majority, stop the applications and wait for the cluster to ↵Simon MacMullen2013-02-271-4/+17
| | | | | | | | | | | | come back.
* | | cluster_cp_modeSimon MacMullen2013-02-181-0/+26
| |/ |/|
* | stable to defaultSimon MacMullen2013-01-241-1/+1
|\ \ | |/
| * Update copyright 2013bug25343Emile Joubert2013-01-231-1/+1
| |
* | Fix a couple of partition-related specs.Simon MacMullen2012-12-101-1/+1
|/
* Explain whySimon MacMullen2012-11-131-0/+4
|