Re: Adding skip scan (including MDAM style range skip scan) to nbtree

Natalya Aksman <natalya@tigerdata.com>

From: Natalya Aksman <natalya@tigerdata.com>

To: Peter Geoghegan <pg@bowt.ie>

Cc: Masahiro Ikeda <ikedamsh@oss.nttdata.com>, Tomas Vondra <tomas@vondra.me>, Masahiro.Ikeda@nttdata.com, pgsql-hackers@lists.postgresql.org, Masao.Fujii@nttdata.com

Date: 2025-09-10T18:59:02Z

Lists: pgsql-hackers

Commits

Same data as JSON: GET /api/v1/messages/:b64id/commits the thread's linked commits as JSON, with link sources. API reference →

nbtree: Always set skipScan flag on rescan.
- 454c046094ab 19 (unreleased) landed
- bee763aea13f 18.0 landed
meson: Build numeric.c with -ftree-vectorize.
- 9016fa7e3bcd 19 (unreleased) cited
Fix "variable not found in subplan target lists" in semijoin de-duplication.
- b8a1bdc458e3 19 (unreleased) cited
Revert "nbtree: Remove useless row compare arg."
- dd2ce3792754 18.0 landed
nbtree: Remove useless row compare arg.
- 54c6ea8c81db 18.0 cited
Prevent premature nbtree array advancement.
- 5f4d98d4f371 18.0 landed
nbtree: tighten up array recheck rules.
- 7e25c9363a82 18.0 landed
Avoid treating nonrequired nbtree keys as required.
- 0f08df406822 18.0 landed
Adjust overstrong nbtree skip array assertion.
- 9d924dbb3710 18.0 landed
Make NULL tuple values always advance skip arrays.
- b75fedcab791 18.0 cited
Avoid extra index searches through preprocessing.
- b3f1a13f22f9 18.0 landed
Improve nbtree skip scan primitive scan scheduling.
- 21a152b37f36 18.0 landed
Further optimize nbtree search scan key comparisons.
- 8a510275dd6b 18.0 landed
Add nbtree skip scan optimization.
- 92fe23d93aa3 18.0 landed
Improve nbtree array primitive scan scheduling.
- 9a2e2a285a14 18.0 landed
nbtree: Make BTMaxItemSize into object-like macro.
- 426ea611171d 18.0 landed
Show index search count in EXPLAIN ANALYZE, take 2.
- 0fbceae841cb 18.0 landed
Make parallel nbtree index scans use an LWLock.
- 67fc4c9fd7fa 18.0 landed
Show index search count in EXPLAIN ANALYZE.
- 5ead85fbc811 18.0 landed
Avoid nbtree parallel scan currPos confusion.
- b5ee4e52026b 18.0 cited
nbtree: Remove useless 'strat' local variable.
- b6558e4f837e 18.0 landed
Normalize nbtree truncated high key array behavior.
- 79fa7b3b1a44 18.0 landed
Refactor handling of nbtree array redundancies.
- b524974106ac 18.0 landed
Fix nbtree pgstats accounting with parallel scans.
- c00c54a9ac1e 18.0 landed
- fb4f5e58af97 17.0 landed
Avoid parallel nbtree index scan hangs with SAOPs.
- d8adfc18bebf 18.0 landed
- a24bffc021d9 17.0 landed
Show Parallel Bitmap Heap Scan worker stats in EXPLAIN ANALYZE
- 5a1e6df3b84c 18.0 cited
Enhance nbtree ScalarArrayOp execution.
- 5bf748b86bc6 17.0 cited
Skip checking of scan keys required for directional scan in B-tree
- e0b1ee17dc3a 17.0 cited
Instead of using a numberOfRequiredKeys count to distinguish required
- 7ccaf13a06b8 8.2.0 cited

Timescaledb implemented multikey skipscan feature for queries like "select
distinct key1, key2 ... from t_indexed_on_key1_key2". It pins key1 to a
found key value (i.e key1=val1)  to skip over distinct values of key2. Then
after values for (key1=va1) are exhausted the next distinct tuple is
searched with (key1>val1).

In short, this implementation can change the scan key structure from
"key1=val1" to "key1>val1" and back, and not just the key comparison value
(i.e. val1).
It means that so->skipScan can get reset from true to false after the next
call to _bt_preprocess_keys.

But after btrescan resets "so->numberOfKeys = 0", so->skipScan is not reset
to "false" in  _bt_preprocess_keys because of this code:
https://github.com/postgres/postgres/blob/9016fa7e3bcde8ae4c3d63c707143af147486a10/src/backend/access/nbtree/nbtpreprocesskeys.c#L1847
After we set "so->numberOfKeys = 0" we quit on line 1847 before we get to
the line 1874 where we do "so->skipScan = (numSkipArrayKeys > 0);"
https://github.com/postgres/postgres/blob/9016fa7e3bcde8ae4c3d63c707143af147486a10/src/backend/access/nbtree/nbtpreprocesskeys.c#L1874

I.e. if btrescan resets  "so->numberOfKeys = 0",  _bt_preprocess_keys quits
before resetting  so->skipScan to false.
It is not an issue when the scan key structure is not changed in amrescan,
and I see that this is an intended usage.
But in case the intended amrescan usage changes in the future, the issue
may come up.

It's not a priority at the moment as we can reset so->skipScan in our
extension.

Thank you,
Natalya Aksman.

On Wed, Sep 10, 2025 at 12:46 PM Peter Geoghegan <pg@bowt.ie> wrote:

> On Wed, Sep 10, 2025 at 9:53 AM Natalya Aksman <natalya@tigerdata.com>
> wrote:
> > Our Timescaledb extension has scenarios changing ">" quals to "=" and
> back on rescan and it breaks when so->Skipscan needs to be reset from true
> to false.
>
> But the amrescan docs say:
>
> "In practice the restart feature is used when a new outer tuple is
> selected by a nested-loop join and so a new key comparison value is
> needed, but the scan key structure remains the same" [1].
>
> I don't understand why it is that our not resetting the so->Skipscan
> flag within btrescan has any particular significance to Timescaledb,
> relative to all of the other fields that are supposed to be set by
> _bt_preprocess_keys. What is the actual failure you see? Is it an
> assertion failure within _bt_readpage/_bt_checkkeys?
>
> Note that btrescan *does* set "so->numberOfKeys = 0", which will make
> the next call to _bt_preprocess_keys (from _bt_first) perform
> preprocessing from scratch. This should set so->Skipscan from scratch
> on each rescan (along with every other field set by preprocessing). It
> seems like that should work for you (in spite of the fact that you're
> doing something that seems at odds with the index AM API).
>
> [1] https://www.postgresql.org/docs/current/index-functions.html
> --
> Peter Geoghegan
>