phorge-phorge

mirror of https://we.phorge.it/source/phorge.git synced 2024-12-02 19:52:44 +01:00

Author	SHA1	Message	Date
Mukunda Modell	e41c25de50	Support multiple fulltext search clusters with 'cluster.search' config Summary: The goal is to make fulltext search back-ends more extensible, configurable and robust. When this is finished it will be possible to have multiple search storage back-ends and potentially multiple instances of each. Individual instances can be configured with roles such as 'read', 'write' which control which hosts will receive writes to the index and which hosts will respond to queries. These two roles make it possible to have any combination of: * read-only * write-only * read-write * disabled This 'roles' mechanism is extensible to add new roles should that be needed in the future. In addition to supporting multiple elasticsearch and mysql search instances, this refactors the connection health monitoring infrastructure from PhabricatorDatabaseHealthRecord and utilizes the same system for monitoring the health of elasticsearch nodes. This will allow Wikimedia's phabricator to be redundant across data centers (mysql already is, elasticsearch should be as well). The real-world use-case I have in mind here is writing to two indexes (two elasticsearch clusters in different data centers) but reading from only one. Then toggling the 'read' property when we want to migrate to the other data center (and when we migrate from elasticsearch 2.x to 5.x) Hopefully this is useful in the upstream as well. Remaining TODO: * test cases * documentation Test Plan: (WARNING) This will most likely require the elasticsearch index to be deleted and re-created due to schema changes. Tested with elasticsearch versions 2.4 and 5.2 using the following config: ```lang=json "cluster.search": [ { "type": "elasticsearch", "hosts": [ { "host": "localhost", "roles": { "read": true, "write": true } } ], "port": 9200, "protocol": "http", "path": "/phabricator", "version": 5 }, { "type": "mysql", "roles": { "write": true } } ] Also deployed the same changes to Wikimedia's production Phabricator instance without any issues whatsoever. ``` Reviewers: epriestley, #blessed_reviewers Reviewed By: epriestley, #blessed_reviewers Subscribers: Korvin, epriestley Tags: #elasticsearch, #clusters, #wikimedia Differential Revision: https://secure.phabricator.com/D17384	2017-03-26 08:16:47 +00:00
epriestley	b2cdebefea	Fix two errors from the error logs Summary: Found these in the `secure` error logs: one bad call, one bad column. Test Plan: Searched for empty string. Double-checked method name. Reviewers: chad Reviewed By: chad Differential Revision: https://secure.phabricator.com/D16948	2016-11-26 07:50:57 -08:00
epriestley	8c89fc38fc	Allow persistent connections to be configured per database host Summary: Ref T11044. Fixes T11672. In T11672, persistent connections seem to work fine, but they can require `max_connections` and other settings to be raised. Since most users don't need them, make them an advanced option. Test Plan: Configured persistent connections, loaded some pages, observed persistent connections get used. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044, T11672 Differential Revision: https://secure.phabricator.com/D16913	2016-11-22 10:55:45 -08:00
epriestley	e6bfa1bd23	Remove "mysql.configuration-provider" configuration option Summary: Ref T11044. This was old Facebook cruft for reading configuration from SMC (and maybe doing some other questionable things). See D183. (See also D175 for discussion of this from 2011.) In modern Phabricator, you can subclass `SiteConfig` to provide dynamic configuration, and we do so in the Phacility cluster. This lets you change any config, and change in response to requests (e.g., for instancing) and is generally more powerful than this mechanism was. This configuration provider theoretically let you roll your own replication or partitioning, but in practice I believe no one ever did, and no one ever could have anyway without more support in the upstream (for migrations, read-after-write, etc). Test Plan: - Grepped for removed option. - Browsed around with clustering off. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044 Differential Revision: https://secure.phabricator.com/D16911	2016-11-22 09:24:46 -08:00
epriestley	4da74166fe	When storage is partitioned, refuse to serve requests unless web and databases agree on partitioning Summary: Ref T11044. One popular tool in a modern operations environment is Puppet. The primary purpose of this tool is to randomly revert hosts to older or different configurations. Introducing an element of chaotic unpredictability into operations trains staff to be on high alert at all times, rather than lulled into complacency by predictability or consistency. When Puppet reverts a Phabricator host's configuration to an older version, we might start writing data to a lot of crazy places where it shouldn't go. This will create a big sticky mess that is virtually impossible to undo, mostly because we'll get two files with ID 123 or two tasks with ID 456 or whatever else and good luck with that. Instead, after changing the partition layout, require `bin/storage partition` to be run. This writes a copy of the config everywhere. Then, when we start serving web requests, make sure every database has the exact same config. This will foil Puppet by refusing to run requests on hosts it has reverted. Test Plan: - Changed partition configuration. - Ran Phabricator. - FOILED! - Ran `bin/storage partition` to sync config. - Things worked again. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044 Differential Revision: https://secure.phabricator.com/D16910	2016-11-22 04:15:46 -08:00
epriestley	bac27fb403	Remove "mysql.implementation" configuration Summary: Ref T11044. Fixes T10931. This option has essentially never been useful for anything, and we've picked the best implementation for a long time (MySQLi if available, MySQL if not). I am not aware of any reason to ever set this manually. If someone comes up with some bizarre but legitimate use case that I haven't thought of, we can modularize it. Test Plan: Browsed around. Grepped for `mysql.implementation`. Reviewers: chad Reviewed By: chad Maniphest Tasks: T10931, T11044 Differential Revision: https://secure.phabricator.com/D16909	2016-11-22 04:15:34 -08:00
epriestley	bcfd515b32	Run all minor setup checks on all configured database hosts Summary: Fixes T10759. Fixes T11817. This runs all the general sanity/configuration checks on all the active servers. None of these warnings are very important, and this doesn't change any logical stuff. Depends on D16904. Test Plan: Painstakingly triggered each warning, verified that they rendered correctly and that messages told me which host was affected. Reviewers: chad Reviewed By: chad Maniphest Tasks: T10759, T11817 Differential Revision: https://secure.phabricator.com/D16905	2016-11-21 15:55:54 -08:00
epriestley	326d5bf800	Detect replicating masters and fatal (also, warn on nonreplicating replicas) Summary: Ref T10759. Check master/replica status during startup. After D16903, this also means that we check this status after a database comes back online after being unreachable. If a master is replicating, fatal (since this can do a million kinds of bad things). If a replica is not replicating, warn (this just means the replica is behind so some data is at risk). Also: if your masters were actually configured properly (mine weren't until this change detected it), we would throw away patches as we applied them, so they would only apply to the //first// master. Instead, properly apply all migration patches to all masters. Test Plan: - Started Phabricator with a replicating master, got a fatal. - Stopped replication on a replica, got a warning. - With two non-replicating masters, upgraded storage. Reviewers: chad Reviewed By: chad Maniphest Tasks: T10759 Differential Revision: https://secure.phabricator.com/D16904	2016-11-21 15:55:22 -08:00
epriestley	55e21565b5	Support application partitioning across multiple masters Summary: Ref T11044. I'm going to hold this until after the release cut, but I think it's good to go. This allows installs to configure multiple masters in `cluster.databases` and partition applications across them (for example, put Maniphest on a dedicated database). When we make a Maniphest connection we go look up which master we should be hitting first, then connect to it. This has at least approximately been planned for many years, so the actual change is largely just making sure that your config makes sense. Test Plan: - Configured `db001.epriestley.com` and `db002.epriestley.com` as master/master. - Partitioned applications between them. - Interacted with various applications, saw writes go to the correct host. - Viewed "Database Servers" and saw partitioning information. - Ran schema upgrades. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044 Differential Revision: https://secure.phabricator.com/D16876	2016-11-19 14:14:39 -08:00
epriestley	558d194302	Update `bin/storage` workflows to accommodate multiple masters Summary: Depends on D16847. Ref T11044. This updates the remaining storage-related workflows from the CLI to accommodate multiple masters. Test Plan: - Configured multiple masters. - Ran all `bin/storage` workflows. - Ran `arc unit --everything`. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044 Differential Revision: https://secure.phabricator.com/D16848	2016-11-12 16:37:47 -08:00
epriestley	bc15eee3f2	Update SchemaQuery and the web UI to accommodate multiple master databases Summary: Depends on D16115. Ref T11044. In the brave new world of multiple masters, we need to check the schemata on each master when looking for missing storage patches, keys, schema changes, etc. This realigns all the "check out what's up with that schema" calls to work for multiple hosts, and updates the web UI to include a "Server" column and allow you to browse per-server. This doesn't update `bin/storage`, so it breaks things on its own (and unit tests probably won't pass). I'll update that in the next change. Test Plan: Configured local environment in cluster mode with multiple masters, saw both hosts' status reported in web UI. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044 Differential Revision: https://secure.phabricator.com/D16847	2016-11-12 16:36:52 -08:00
epriestley	ecc598f18d	Support multiple database masters and convert easy callers Summary: Ref T11044. This moves toward partitioned application databases: - You can define multiple masters. - Convert all the easily-convertible code to become multi-master aware. This doesn't convert most of `bin/storage` or "Config > Database (Stuff)" yet, as both are quite involved. They still work for now, but only operate on the first master instead of all masters. Test Plan: Configured multiple masters, browsed around, ran `bin/storage` commands, ran `bin/storage --host ...`. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11044 Differential Revision: https://secure.phabricator.com/D16115	2016-11-12 16:30:20 -08:00
epriestley	d0013d0898	Distinguish between unreachable cluster database hosts and missing MySQL databases Summary: Fixes T11577. When we connect to a host and try to select a database which does not exist, we currently treat it as though the host wasn't reachable. This isn't correct, and prevents storage from being initialized while already in cluster mode, since the "config" database won't exist yet the first time we connect. Instead, distinguish between `AphrontSchemaQueryException` (thrown on connection if the requested database is not present) and other errors. Test Plan: - Put Phabricator into cluster database mode (`cluster.databases = ...`). - Swapped `storage.default-namespace` to force initialization of a new install. - Ran `bin/storage upgrade`. - Before patch: Immediate fatal about unreachablility. - After patch: Database initialized. - Also ran initialization steps in tranditional single-host mode (`cluster.databases` empty, `mysql.host` configured). Reviewers: chad Reviewed By: chad Maniphest Tasks: T11577 Differential Revision: https://secure.phabricator.com/D16489	2016-09-02 08:23:21 -07:00
epriestley	01040e4573	Correctly disinguish between "0 seconds behind master" and "not replicating" Summary: Fixes T11159. We get two different values here (`NULL` and `0`) with different meanings. Test Plan: - Ran `STOP SLAVE;`. - Saw this: {F1710181} - Ran `START SLAVE;`. - Back to normal. Reviewers: chad Reviewed By: chad Maniphest Tasks: T11159 Differential Revision: https://secure.phabricator.com/D16225	2016-07-03 18:14:07 -07:00
epriestley	ac35246d0d	Never sever non-cluster database; write more read-only documentation Summary: Ref T4571. Write more of the missing documentation sections and clarify a few things. Since the "replicating master" check needs a special permission, imposes a performance penalty, is probably very difficult to misconfigure, and likely not a big deal anyway, just drop the idea of trying to automatically detect + prevent it. We still show if it's an issue on the status page, provided we have permission to check. When you don't have any cluster databases configured, never stop trying to connect to the default master database. We might want to do this eventually as load reduction, but just don't muddy the waters too much for now while things stabilize. Test Plan: - Tested functionality in cluster, non-cluster, and degraded-cluster modes. - Used status console to monitor a health check cycle. - Read docs. Reviewers: chad Reviewed By: chad Maniphest Tasks: T4571 Differential Revision: https://secure.phabricator.com/D15679	2016-04-11 08:44:11 -07:00
epriestley	ebff07d019	Automatically sever databases after prolonged unreachability Summary: Ref T4571. When a database goes down briefly, we fall back to replicas. However, this fallback is slow (not good for users) and keeps sending a lot of traffic to the master (might be bad if the root cause is load-related). Keep track of recent connections and fully degrade into "severed" mode if we see a sequence of failures over a reasonable period of time. In this mode, we send much less traffic to the master (faster for users; less load for the database). We do send a little bit of traffic still, and if the master recovers we'll recover back into normal mode seeing several connections in a row succeed. This is similar to what most load balancers do when pulling web servers in and out of pools. For now, the specific numbers are: - We do at most one health check every 3 seconds. - If 5 checks in a row fail or succeed, we sever or un-sever the database (so it takes about 15 seconds to switch modes). - If the database is currently marked unhealthy, we reduce timeouts and retries when connecting to it. Test Plan: - Configured a bad `master`. - Browsed around for a bit, initially saw "unrechable master" errors. - After about 15 seconds, saw "major interruption" errors instead. - Fixed the config for `master`. - Browsed around for a while longer. - After about 15 seconds, things recovered. - Used "Cluster Databases" console to keep an eye on health checks: it now shows how many recent health checks were good: {F1213397} Reviewers: chad Reviewed By: chad Maniphest Tasks: T4571 Differential Revision: https://secure.phabricator.com/D15677	2016-04-11 08:43:52 -07:00
epriestley	146fb646f9	Automatically degrade to read-only mode when unable to connect to the master Summary: Ref T4571. If we fail to connect to the master, automatically try to degrade into a temporary read-only mode ("UNREACHABLE") for the remainder of the request, if possible. If the request was something like "load the homepage", that'll work fine. If it was something like "submit a comment", there's nothing we can do and we just have to fail. Detecting this condition imposes a performance penalty: every request checks the connection and gives the database a long time to respond, since we don't want to drop writes unless we have to. So the degraded mode works, but it's really slow, and may perpetuate the problem if the root issue is load-related. This lays the groundwork for improving this case by degrading futher into a "SEVERED" mode which will persist across requests. In the future, if several requests in a short period of time fail, we'll sever the database host and refuse to try to connect to it for a little while, connecting directly to replicas instead (basically, we're "health checking" the master, like a load balancer would health check a web application server). This will give us a better (much faster) degraded mode in a major service disruption, and reduce load on the master if the root cause is load-related, giving it a better chance of recovering on its own. Test Plan: - Disabled master in config by changing the host/username, got degraded automatically to UNREACAHBLE mode immediately. - Faked full SEVERED mode, requests hit replicas and put me in the mode properly. - Made stuff work, hit some good pages. - Hit some non-cluster pages. Reviewers: chad Reviewed By: chad Maniphest Tasks: T4571 Differential Revision: https://secure.phabricator.com/D15674	2016-04-10 12:20:13 -07:00
epriestley	e0a8cac703	When no master database is configured, automatically degrade to read-only mode Summary: Ref T4571. If `cluster.databases` is configured but only has replicas, implicitly drop to read-only mode and send writes to a replica. Test Plan: - Disabled the `master`, saw Phabricator automatically degrade into read-only mode against replicas. - (Also tested: explicit read-only mode, non-cluster mode, properly configured cluster mode). Reviewers: chad Reviewed By: chad Maniphest Tasks: T4571 Differential Revision: https://secure.phabricator.com/D15672	2016-04-10 12:19:55 -07:00
epriestley	c178f29cdb	Use new first-class MySQL timeout support in Phabricator Summary: Fixes T6710. After D15669, we support a proper timeout parameter, so we don't need this hack anymore. Test Plan: See D15669: forced a MySQL connector, set a low timeout, set a bad database, saw fast failures. Reviewers: chad Reviewed By: chad Maniphest Tasks: T6710 Differential Revision: https://secure.phabricator.com/D15670	2016-04-10 12:19:00 -07:00
epriestley	6a4a9bb2d2	When `cluster.databases` is configured, read the master connection from it Summary: Ref T4571. Ref T10759. Ref T10758. This isn't complete, but gets most of the job done: - When `cluster.databases` is set up, most things ignore `mysql.host` now. - You can `bin/storage upgrade` and stuff works. - You can browse around in the web UI and stuff works. There's still a lot of weird tricky stuff to navigate, and this has real no advantages over configuring a single server yet (no automatic failover, etc). Test Plan: - Configured `cluster.databases` to point at my `t1.micro` hosts in EC2 (master + replica). - Ran `bin/storage upgrade`, got a new install setup on them properly. - Survived setup warnings, browsed around. - Switched back to local config, ran `bin/storage upgrade`, browsed around, went through setup checks. - Intentionally broke config (bad hosts, no masters) and things seemed to react reasonably well. Reviewers: chad Reviewed By: chad Maniphest Tasks: T4571, T10758, T10759 Differential Revision: https://secure.phabricator.com/D15668	2016-04-10 12:18:42 -07:00
epriestley	0439645d5b	Add a "Database Cluster Status" console in Config Summary: Ref T4571. The configuration option still doesn't do anything, but add a status panel for basic setup monitoring. Test Plan: Here's what a good version looks like: {F1212291} Also faked most of the errors it can detect and got helpful diagnostic messages like this: {F1212292} Reviewers: chad Reviewed By: chad Maniphest Tasks: T4571 Differential Revision: https://secure.phabricator.com/D15667	2016-04-09 20:34:13 -07:00

21 commits