Skip to content

index not updated correctly when removing nodes from placement #4196

Description

@BertHartm

When scaling down a cluster, we're noticing that some data becomes unavailable for query. It appears in the form of partial results when querying for older data (from before the scale down).

We're also noticing that database_tick_index_num_docs remains flat for each node through the scale down, and then jumps up once the node is restarted. The effect if summing across all nodes the cluster is that the metric drops (when the old node is removed), and recovers to prior level when the remaining nodes restart.

General Issues

What service is experiencing the issue? (M3Coordinator, M3DB, M3Aggregator, etc)

m3db

What is the configuration of the service? Please include any YAML files, as well as namespace / placement configuration (with any sensitive information anonymized if necessary).

can provide if required, but I think this might be general
RF=3

How are you using the service? For example, are you performing read/writes to the service via Prometheus, or are you using a custom script?

issue relates to reads happening via remote read

Is there a reliable way to reproduce the behavior? If so, please provide detailed instructions.

It appears to be consistent when removing nodes from placements. It's more obvious when the clusters are small as more of the index is affected.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions