Intermittent DateRangeIncludingNowQuery shard failures on Fess 15.8 / OpenSearch 3.8

Environment

- Fess: 15.8.0
- OpenSearch: 3.8.0
- Docker images:
  - `ghcr.io/codelibs/fess:15.8.0`
  - `ghcr.io/codelibs/fess-opensearch:3.8.0`

Problem

Fess occasionally displays the warning:

> The search processing time exceeded the limit. The displayed results may be incomplete.

However, the actual search is fast and OpenSearch reports:

"timed_out": false

For example, a search completed in about 150 ms.

The Fess log records the request as:

[SEARCH TIMEOUT]

but the OpenSearch response contains failed shards with:

unsupported_operation_exception
Query DateRangeIncludingNowQuery(...) does not implement createWeight

Relevant Fess facet configuration

The default date facets are enabled:

query.facet.queries=\
labels.facet_timestamp_title:\
labels.facet_timestamp_1day=timestamp:[now/d-1d TO *]\t\
labels.facet_timestamp_1week=timestamp:[now/d-7d TO *]\t\
labels.facet_timestamp_1month=timestamp:[now/d-1M TO *]\t\
labels.facet_timestamp_1year=timestamp:[now/d-1y TO *]\n\

Reproduction directly against OpenSearch

The problem can also be reproduced directly against the Fess OpenSearch index, without using the Fess UI.

Example request:

POST fess.search/_search
{
  "size": 0,
  "query": {
    "match_all": {}
  },
  "aggs": {
    "day": {
      "filter": {
        "range": {
          "timestamp": {
            "gte": "now-1d/d"
          }
        }
      }
    },
    "week": {
      "filter": {
        "range": {
          "timestamp": {
            "gte": "now-7d/d"
          }
        }
      }
    },
    "month": {
      "filter": {
        "range": {
          "timestamp": {
            "gte": "now-1M/d"
          }
        }
      }
    },
    "year": {
      "filter": {
        "range": {
          "timestamp": {
            "gte": "now-1y/d"
          }
        }
      }
    }
  }
}

The failure is intermittent.

An example failed response is:

"_shards": {
  "total": 5,
  "successful": 4,
  "failed": 1
}

with:

unsupported_operation_exception:
Query DateRangeIncludingNowQuery(IndexOrDocValuesQuery(...))
does not implement createWeight

Other identical requests return:

"_shards": {
  "total": 5,
  "successful": 5,
  "failed": 0
}

Tests performed

Changing the date math syntax from:

now/d-1d

to:

now-1d/d

does not solve the issue.

In one test series with now-based ranges, 4 out of 20 identical requests had one failed shard.

As a control test, the same aggregation was executed repeatedly with fixed absolute timestamps instead of now.

Approximately 100 repeated requests using fixed dates completed successfully with:

successful=5
failed=0

This strongly suggests that the problem is related to DateRangeIncludingNowQuery.

Expected behavior

Relative date facets using now should not cause intermittent shard failures.

Additionally, Fess should probably not report this condition as:

[SEARCH TIMEOUT]

when OpenSearch reports:

"timed_out": false

The search itself is fast. The problem appears to be a partial shard failure in the date facet aggregation.

Questions

  1. Is this a known issue with OpenSearch 3.8.0?

  2. Is there a recommended workaround for Fess 15.8.0?

  3. Should Fess distinguish partial shard failures from actual query timeouts in the user-facing warning?

  4. If this is confirmed as a bug, should I create a GitHub issue for Fess, or is this better reported directly to OpenSearch?

Thanks for the detailed report and the reproduction. They made this easy to track down.

  1. Is this a known issue?

It was not a known issue, so we have reported it to OpenSearch: [BUG] filter aggregation with `now` date math intermittently fails with "DateRangeIncludingNowQuery ... does not implement createWeight" under concurrent segment search · Issue #23104 · opensearch-project/OpenSearch · GitHub
Our investigation shows that it is an OpenSearch bug introduced in 3.2.0. It is not caused by Fess or by the date math: now/d-1d, now-1d/d and now-1d all fail the same way. We could reproduce it on plain OpenSearch 3.2.0 and 3.8.0, but not on 3.1.0. It only happens when a single-bucket filter aggregation uses a now-based range and concurrent segment search is active. Fess builds its date facets exactly that way.

  1. Workaround

Disable concurrent segment search on the Fess index. fess.search is an alias, so first look up the actual index name:

GET _cat/aliases/fess.search?v

Then apply the setting to that index, for example fess.20260917:

PUT fess.20260917/_settings
{ "index.search.concurrent_segment_search.mode": "none" }

This is a dynamic setting, so it takes effect immediately; you do not need to close the index or restart OpenSearch. In our tests the failures stopped completely with none. The only effect is that aggregations no longer run in parallel within a shard. If the Fess index is ever recreated, for example by a reindex, apply the setting again to the new index.

  1. Timeout vs. shard failure message

You are right. Fess 15.8 treats any failed shard as a “timeout”. This is fixed in the upcoming Fess 15.9 (fix(search): tell a query timeout apart from a failed shard by marevol · Pull Request #3440 · codelibs/fess · GitHub): a timeout and a shard failure are now reported separately. Until then, if you see this message on 15.8, the cause is more likely a shard failure than a timeout. The real reason is in the _shards.failures part of the [SEARCH TIMEOUT] log line.

  1. Where to report

No Fess issue is needed. The OpenSearch issue above covers it.

1 Like