RavenDB supports two search engines for indexes: Corax and Lucene. The selected engine determines how RavenDB builds and queries an index. The choice of engine is made when creating a database or index, and while it can be changed later, doing so triggers a full index rebuild.
The ServiceControl error and audit databases have a specific workload: messages are ingested continuously, expired messages are deleted continuously by the retention process, and the data is queried only occasionally, when ServicePulse is used. Load testing of this workload showed that, for the index definitions ServiceControl uses, Lucene indexes:
- take up less storage
- use less memory
- keep up with ingestion and deletion with less index lag
- query faster
Meanwhile, Corax demonstrated more performance and stability issues on large ServiceControl databases. Because of this, starting with ServiceControl version 6.20, new databases default to using Lucene, and migrating existing databases to Lucene is recommended.
Default search engine per version
| ServiceControl version | Search engine for new databases | Existing databases |
|---|---|---|
| 5.x and 6.0–6.19 | Corax | Keep the engine they were created with |
| 6.20 and later | Lucene | Keep the engine they were created with |
As the table above illustrates, upgrading ServiceControl does not impact the search engine of existing databases. This is because changing the engine triggers a full rebuild of every index which, depending on the available computing power, can take days on very large databases. In addition, while the rebuild runs, the ingestion and indexing rates are degraded. Thus, migration should be planned and scheduled for each environment. See Should existing indexes be migrated? and Migrating existing indexes to Lucene for more details.
Monitoring instances do not use RavenDB and are not affected.
Detecting indexes that use Corax
Available in version 6.20
Error and audit instances report indexes using Corax in two ways:
- A custom check named
Error Database Search Engine(error instance) orAudit Database Search Engine(audit instance), visible in ServicePulse. The check fails when at least one index uses Corax. It is evaluated hourly. - A warning in the instance log at every start-up.
Both list the affected indexes and contain the following message:
The following RavenDB index(es) use the Corax search engine:
. Lucene indexes are smaller, use less memory and perform better for ServiceControl workloads, and are the default for new databases. Consider switching these indexes to Lucene. Note that switching triggers a full rebuild of the index: on very large databases this can take days depending on the available compute, and while the rebuild is running ingestion and indexing rates can be degraded. Plan the switch accordingly.<database>/ <index>
The search engine of an index can also be inspected in RavenDB Studio, on the Configuration tab of the index, or in the Indexes list where each index shows its engine.
Should existing indexes be migrated?
Yes. Migrating existing error and audit databases to Lucene is recommended for all instances. Corax has shown performance and stability issues on large ServiceControl databases, and these issues get worse as the database grows. Migrating proactively, while the database is small and the instance is healthy, keeps the rebuild short and efficient. This avoids a more complicated migration when the instance is already struggling.
Migrate as soon as possible when the instance shows one or more of the following symptoms. They indicate that the Corax indexes can no longer keep up with the load:
- Frequent or persistent index lag, reported by the stale indexes custom check
- High RAM utilization or RavenDB dirty memory warnings
- High CPU utilization caused by indexing
- Corrupted indexes or a lengthy database recovery after a service shutdown (see Audit instances: Corrupted indexes or corrupted database after a service shutdown for troubleshooting this problem)
- Database storage growth that is dominated by the size of the indexes
Instances without these symptoms should be migrated in the next planned maintenance window. Because the rebuild requires downtime or degraded ingestion, plan the migration separately for each environment. Migrate development and test instances first to estimate the rebuild duration for production. The Error Database Search Engine or Audit Database Search Engine custom check continues to fail until all indexes use Lucene.
Before migrating, back up the database and estimate the rebuild duration. The rebuild has to process every document in the database. Extrapolate from a smaller instance, or from the time the last database upgrade took, and schedule the migration in a maintenance window. Consider temporarily adding CPU and RAM to the host until the rebuild completes.
Migrating existing indexes to Lucene
The migration is performed per index in RavenDB Studio.
On ServiceControl versions before 6.20, the migrated index must also be locked afterward. These versions recreate their index definitions at every start-up and will reset the index to the database default (Corax) and trigger another rebuild otherwise.
The indexes with the highest load, and therefore the ones that benefit most, are:
| Instance | Index |
|---|---|
| Error | FailedMessageViewIndex, MessagesViewIndex |
| Audit | MessagesViewIndex or MessagesViewIndexWithFullTextSearch (depending on whether full-text search on message bodies is enabled) |
Other indexes can be migrated using the same procedure. Migrate one index at a time and wait for it to become non-stale before migrating the next one to limit the impact on ingestion.
1. Access the RavenDB Studio
- Windows deployment: Start the instance in maintenance mode and click Launch RavenDB Studio.
- Container deployment: Stop the ServiceControl container and open RavenDB Studio on port
8080of the database container. - External RavenDB server: Open RavenDB Studio for the server that hosts the ServiceControl database.
Running the migration while the instance is stopped (maintenance mode) is recommended. It avoids ingestion competing with the rebuild for CPU and I/O and prevents the instance from resetting the index before it has been locked. Messages will accumulate in the error and audit queues while the instance is stopped; ensure the queues have enough capacity for the expected duration.
2. Change the search engine of the index
- In RavenDB Studio, select the ServiceControl database and open Indexes > List of Indexes.
- Click the index to edit it.
- Open the Configuration tab.
- Change Search engine from
CoraxorCorax (inherited)toLucene. - Click Save.
RavenDB creates a replacement index that uses Lucene and runs side-by-side with the existing Corax index. The existing index will keep serving queries until the replacement has caught up.
3. Swap the indexes
In the List of Indexes, the index shows the replacement being built. Once the replacement is no longer stale, RavenDB swaps it in automatically and deletes the Corax index. RavenDB Studio also offers swap now:
- Swapping immediately frees the storage of the Corax index right away but queries return stale results until the Lucene index has been fully rebuilt.
- Waiting for the automatic swap keeps queries accurate but temporarily requires storage for both indexes.
Under constant ingestion the replacement index may never be reported as non-stale and the automatic swap may never happen. This is another reason to perform the migration while the instance is in maintenance mode, or to swap the indexes manually.
4. Lock the index
Locking the index is required to keep a non-default index configuration on ServiceControl versions prior to 6.20.0. On version 6.20.0 and later, this step can be skipped.
While still in RavenDB Studio, click the 🔓 Unlocked button of the migrated index and change it to 🔒 Locked (ignore) (lock modes). RavenDB Studio confirms with Lock mode was set to: Locked (ignore).
A locked index is left untouched when ServiceControl recreates its index definitions at start-up, so the index stays on Lucene.
Locking an index also means that index definition changes shipped with future ServiceControl versions are not applied to it. Check the upgrade guide of each new version for changes to the locked indexes; if an index definition changes, unlock the index, let ServiceControl update it, and repeat this migration for it.
5. Restart the instance
Stop maintenance mode or start the ServiceControl container. The next start-up no longer logs a warning for the migrated index, and the Error Database Search Engine / Audit Database Search Engine custom check passes once all indexes use Lucene.
Migrating the whole database
Instead of migrating indexes one by one, the database default can be changed so that all indexes, including future ones, use Lucene without locking:
- In RavenDB Studio, open Settings > Database Settings for the ServiceControl database.
- Set both
Indexing.andStatic. SearchEngineType Indexing.toAuto. SearchEngineType Luceneand save. - Reload the database when prompted.
- Reset each index (List of Indexes > index menu > Reset) so that it is rebuilt with the new engine.
Resetting an index deletes it and rebuilds it from scratch; queries against it return stale results until the rebuild completes. Because all indexes are rebuilt, this approach causes a longer period of degraded performance than migrating individual indexes, but does not require indexes to be locked and applies to indexes added by future ServiceControl versions as well. Because the custom check only passes once all indexes use Lucene, this approach is the most direct way to complete the migration.