Proposal and Design for Root-Level Nested Aggregation Fields #1363
Replies: 3 comments 10 replies
|
I'm excited to see this proposal--I think it's going to be a big unlock. I also see it as merely being part 1 in a larger story of a direction I'd like ElasticGraph to go: I'd like ElasticGraph to treat nested documents as first-class documents--besides supporting direct aggregation of them (this proposal), I'd also like to see us add support for querying raw nested documents. For example, see the warning in our list filtering docs--in a case like that music schema, it would be useful to be able to directly filter and paginate I bring this up because it's good to bear in mind as we design this feature: the design choices we make for this feature should work as-is in the future with direct nested document queries. (For example: the GraphQL input shape for how you filter on parent document fields should be the same for the two cases). This proposal has several points where there are interesting decisions to make. I'm going to split off separate responses to discuss each so that we can have threaded discussions about each part of the design. |
Ancestor filteringTo me, this is one of the main open questions and I'd like to enumerate the options carefully and consider them all. To start with, I'm going to use the music schema used on EG's public website and its bootstrapping template. (I find it easier to work with than the example you gave: the
For a root field aggregating tracks (
Also relevant: filtering on root document fields is essential for efficiency (index targeting with rollover, shard routing), so the root level's filter needs to be cleanly extractable and prominent, not an afterthought. With that, the shapes I see: A. Outside-in, path-nested, target fields inline (as proposed): # SDL: a new chain of input types per exposed path. Every level's fields are
# duplicated onto a path-specific type because each level mixes in a key that
# doesn't exist on the standard FilterInput for that type.
input ArtistAlbumTrackAggregationsFilterInput {
# ...all TrackFilterInput fields, duplicated inline (anyOf omitted or runtime-validated)...
lengthInSeconds: IntFilterInput
artist: ArtistAlbumTrackAggregationsArtistFilterInput
}
input ArtistAlbumTrackAggregationsArtistFilterInput {
# ...all ArtistFilterInput fields, duplicated, except `albums` changes meaning/type:
bio: ArtistBioFilterInput
albums: ArtistAlbumTrackAggregationsAlbumFilterInput # containing-album semantics, no anySatisfy
}
input ArtistAlbumTrackAggregationsAlbumFilterInput {
# ...all AlbumFilterInput fields, duplicated...
name: StringFilterInput
}
artistAlbumTrackAggregations(filter: ArtistAlbumTrackAggregationsFilterInput, ...): ...artistAlbumTrackAggregations(filter: {
artist: {
bio: {yearFormed: {gt: 2000}}
albums: {name: {equalToAnyOf: ["Greatest Hits"]}} # the *containing* album
}
lengthInSeconds: {gt: 500} # Track fields inline
})B. Inside-out parent chain, target fields inline: # SDL: a new type per level except the root, which can reuse ArtistFilterInput
# as-is (it's the end of the chain--no extra key needed on it).
input ArtistAlbumTrackAggregationsFilterInput {
# ...all TrackFilterInput fields, duplicated inline...
lengthInSeconds: IntFilterInput
parentAlbum: AlbumAsParentOfTrackFilterInput
}
input AlbumAsParentOfTrackFilterInput {
# ...all AlbumFilterInput fields, duplicated...
name: StringFilterInput
parentArtist: ArtistFilterInput # reused!
}
artistAlbumTrackAggregations(filter: ArtistAlbumTrackAggregationsFilterInput, ...): ...filter: {
lengthInSeconds: {gt: 500}
parentAlbum: {
name: {equalToAnyOf: ["Greatest Hits"]}
parentArtist: {bio: {yearFormed: {gt: 2000}}}
}
}C. Separate argument per level: # SDL: no new input types at all--every argument reuses an existing FilterInput.
artistAlbumTrackAggregations(
filter: TrackFilterInput
albumFilter: AlbumFilterInput
artistFilter: ArtistFilterInput
first: Int
after: Cursor
): ...artistAlbumTrackAggregations(
filter: {lengthInSeconds: {gt: 500}} # TrackFilterInput, unchanged meaning
albumFilter: {name: {equalToAnyOf: ["Greatest Hits"]}} # AlbumFilterInput: the containing album
artistFilter: {bio: {yearFormed: {gt: 2000}}} # ArtistFilterInput: the root document
)D. Single # SDL: one thin wrapper type per exposed path; every value reuses an existing
# FilterInput. Deliberately no anyOf/not on the wrapper (see observation 1).
# Named after the nested path, not the aggregation field, so the same type can
# be shared with the future raw-nested-document query field for this path.
input NestedArtistAlbumTrackFilterInput {
track: TrackFilterInput
album: AlbumFilterInput
artist: ArtistFilterInput
}
artistAlbumTrackAggregations(filter: NestedArtistAlbumTrackFilterInput, ...): ...filter: {
track: {lengthInSeconds: {gt: 500}}
album: {name: {equalToAnyOf: ["Greatest Hits"]}}
artist: {bio: {yearFormed: {gt: 2000}}}
}E. Hybrid of C and D: target # SDL: one thin wrapper per exposed path covering only the ancestor levels; the
# target keeps the standard filter: convention. No anyOf/not on the wrapper.
# Named after the path so the future raw-nested-document query field can share it.
input NestedArtistAlbumTrackParentFiltersInput {
album: AlbumFilterInput
artist: ArtistFilterInput
}
artistAlbumTrackAggregations(
filter: TrackFilterInput
parentFilters: NestedArtistAlbumTrackParentFiltersInput
...
): ...artistAlbumTrackAggregations(
filter: {lengthInSeconds: {gt: 500}}
parentFilters: {
album: {name: {equalToAnyOf: ["Greatest Hits"]}}
artist: {bio: {yearFormed: {gt: 2000}}}
}
)F. Global # SDL: every nested type's *existing* FilterInput gains a reserved `parent` key,
# with one sub-key per possible containing type. New types are per-*type*
# (not per-path), and today's sub_aggregations filters would gain ancestor
# filtering for free--no new root-field-specific filter machinery at all.
input TrackFilterInput {
# ...existing fields, plus:
parent: NestedParentOfTrackFilterInput
}
input NestedParentOfTrackFilterInput {
album: AlbumFilterInput # which itself now has parent: {artist: ...}
}
artistAlbumTrackAggregations(filter: TrackFilterInput, ...): ...filter: {
lengthInSeconds: {gt: 500}
parent: {album: {
name: {equalToAnyOf: ["Greatest Hits"]}
parent: {artist: {bio: {yearFormed: {gt: 2000}}}}
}}
}How they compare:
A couple of other notes:
I'd rule out A and B, but I see genuine strengths in each of the others:
On the proposed shape (A) specifically, my concerns are: it needs a new chain of generated types anyway, it has the target-field collision problem (nothing stops |
|
Stepping back from the filtering details in my other comment, I want to float an alternative to the entire proposal: instead of adding new root-level aggregation fields, what if we made the existing aggregations API able to express "ungrouped"? The query this proposal enables can almost be written today: artistAggregations(filter: {bio: {yearFormed: {gt: 2000}}}) { # root filter: index/shard targeting works as usual
nodes { # no groupedBy selected--one implicit bucket
subAggregations {
albums(filter: {name: {equalToAnyOf: ["Greatest Hits"]}}) { # ancestor filter, "containing album" semantics
nodes { # again, one implicit bucket
subAggregations {
tracks(filter: {lengthInSeconds: {gt: 500}}, first: 10) {
nodes { groupedBy { ... } aggregatedValues { ... } }
}
}
}
}
}
}
}Ancestor filtering already exists in this shape--each level's existing Idea: an type ArtistAggregationConnection {
# ...existing nodes / edges / pageInfo...
ungrouped: ArtistUngroupedAggregation # NEW: the single implicit bucket, directly
}
type ArtistUngroupedAggregation {
count: JsonSafeLong
aggregatedValues: ArtistAggregatedValues
subAggregations: ArtistUngroupedSubAggregations
# Note that this has on `groupedBy` since it's ungrouped!
}
type ArtistUngroupedSubAggregations {
# cursor pagination available here, because the types guarantee no grouping above:
albums(filter: AlbumFilterInput, first: Int, after: Cursor): ArtistAlbumUngroupedSubAggregationConnection
}
type ArtistAlbumUngroupedSubAggregationConnection {
pageInfo: PageInfo # real cursors (composite agg)
nodes: [ArtistAlbumSubAggregation!]! # group at this level...
ungrouped: ArtistAlbumUngroupedAggregation # ...or keep drilling down
}artistAggregations(filter: {bio: {yearFormed: {gt: 2000}}}) {
ungrouped {
subAggregations {
albums(filter: {name: {equalToAnyOf: ["Greatest Hits"]}}) {
ungrouped {
subAggregations {
tracks(filter: {lengthInSeconds: {gt: 500}}, first: 10, after: "...") {
pageInfo { hasNextPage endCursor }
nodes { groupedBy { ... } aggregatedValues { ... } }
}
}
}
}
}
}
}What it buys:
What it costs:
These aren't mutually exclusive: we could ship |
Uh oh!
There was an error while loading. Please reload this page.
Problem
sub_aggregationslet a client aggregate nested collections (e.g.Team.seasons_nested), but only as a child of a parent-level aggregation (team_aggregations { nodes { sub_aggregations { seasons_nested { ... } } } }).No real pagination. Sub-aggregation connections expose
firstandhas_next_pagebut noafter/cursors. They're built on a plain terms aggregation with asizecap, not acompositeaggregation. Composite's cursor-based pagination only works today for top-level aggregations, where there's a single grouping level with nothing above it. Once a nested aggregation sits under a parent grouping, a cursor per parent bucket isn't well-defined, so cursor pagination isn't offered at all. This is the core problem this design addresses.Proposed feature
Offer nested types as root-level aggregation query fields:
The query schema would guarantee no grouping happens above the target nested level for this field because it aggregates directly over the nested collection with no parent-level bucketing in the way. That's what makes it valid to use a real
compositeaggregation (nested inside thenestedwrapper), givingafter-based cursor pagination consistent with top-level aggregation.This is additive in that existing
sub_aggregationsare untouched and keep today'sfirst-only behavior. Root-level fields are a new, separate entry point into the same nested collections.As a secondary benefit, these new root-level fields also give clients a shorter path to a deeply-nested collection. To aggregate
Playerrows nested three levels down (Team→TeamSeason→Player), a client using today'ssub_aggregationsmust traverse an extranodes { sub_aggregations { ... } } }layer per intermediate level, even when they have no interest in grouping at Team or Season at all (an intermediate level can be left ungrouped, but itsnodes/sub_aggregationswrapper still has to be written out); a root field skips straight to the target level. This is a convenience; the query already works today, just more verbosely.Ancestor filtering
For a root field aggregating
Playerthree levels deep —Team --(seasons_nested)--> TeamSeason --(some_middle_nested)--> Middle --(players_nested)--> Player, (Middledoesn't exist in theteamsschema today but illustrates a deep nesting case) clients may want to filter on ancestor fields (Team, TeamSeason, Middle), not just Player fields.Proposal: the field's own type (
Player) filters are specified inline infilter, exactly as it would for asub_aggregationsfield's filter today. Every ancestor, including the root, gets its own key, one per step in the path, nested to mirror the actual path structure:teamis the outermost/topmost ancestor key (the root document type);seasons_nestedis the fieldTeamuses to reachTeamSeason, andsome_middle_nestedis the fieldTeamSeasonuses to reachMiddle— each nests inside the previous ancestor's block since each ancestor is only reachable through the one before it.The path from root to target nests one level at a time, each key being the field-path segment used to reach that level; only the target type's own fields are inline. This avoids two problems with a flatter alternative: field name collisions between the root type and the target type (both being flattened into the same object could produce duplicate GraphQL field names), and inconsistent treatment of the root ancestor vs. every other ancestor.
Keying ancestor blocks by field-path segment (rather than type name) keeps them unambiguous even when two ancestor levels share a type, or the same type is reachable via two differently-named paths. The
teamsschema'snested_fields/nested_fields2is exactly this case — each duplicate path already has a distinct field name to key off of. These key names should be configurable via the schema DSL, consistent with how other schema element names are already customizable.Naming / schema surface
Every distinct nested path to a given type needs its own uniquely-named root field. A naming scheme derived automatically from type names alone has potential problems: the
teamsschema'snested_fields/nested_fields2(two distinct paths to the sameTeamNestedFieldstype) would collide if the root field name were built purely from the target type.Proposal: root-level aggregation fields are specified in the schema. A schema author explicitly opts-in by declaring which nested paths get promoted to a root field, and supplies the root field name in doing so, rather than any name being auto-derived from the path or type name. Since names are author-supplied, schema-build-time validation should reject duplicate root field names across different declarations, so a naming mistake surfaces immediately as a build error.
Proposed DSL shape — a declaration on the innermost nested field, naming the root field it should be promoted to:
Known gap: no client-specified sorting (acknowledged, not addressed here)
Neither top-level aggregations nor sub-aggregations support client-specified sorting today — ordering is implicit, driven by grouping keys (ascending, via composite
sourcesfor top-level;doc_countdescending, then re-sorted client-side for stability, for sub-aggregations). This is a documented design choice (config/site/query-api/aggregations.md), not an oversight.Why it matters here: pagination is often only useful paired with sorting — e.g. "top 5 seasons by win count" needs sorting by an
aggregated_valuesmetric, not by grouping key. Adding cursorpagination to root-level fields doesn't make that possible on its own; it only lets you page through buckets in grouping-key order.
Sorting by a metric would need an ES
bucket_sortpipeline aggregation, which interacts awkwardly with composite'safter_keytraversal (which assumes ascending key order) — combining the two isn't a drop-in change and would need its own design.Decision: sorting is out of scope for this design. Root-level aggregation fields ship without client-specified sorting, matching current behavior everywhere else.
Non-goals / deferred
Open questions
Appendix: other aspects considered
reverse_nestedisn't needed (for the base feature)It seemed like ancestor filtering (e.g. filtering a
Playerroot field by aTeamSeasonorTeamfield) would need Elasticsearch'sreverse_nestedaggregation, since the query "starts" at the nested level. It doesn't —reverse_nestedisn't used anywhere in the codebase today, and it isn't required here either. The root field is a resolver-level convenience; the actual ES request still starts at the real index root and nests downward as usual:This is the same filter-composition pattern
Aggregation::Query#filter_detailalready builds for today's parent-filtered sub-aggregations. The only new part is that the resolver skips exposing the intermediate Team/Season grouping levels and returns the innermost composite result directly.reverse_nestedwould only be needed for a different class of query we're not targeting — e.g. "count distinct Teams that have a Player matching some condition," or filtering by a sibling nested collection at the same depth. This is not in scope.Recursive/self-referential nested types: not a concern
A type like
Categorywith a self-referential nested field (subcategories: [Category!]!, allowing arbitrary depth) would need special handling for this feature, but ElasticGraph already disallows that pattern at schema-build time regardless — it rejects any cycle among plain/nested/embedded fields (schema_definition/results.rb,check_for_circular_dependencies!). The only legal self-reference is throughrelates_to_one/relates_to_manyrelationship fields, which are a different mechanism and are already excluded from sub-aggregation path generation.Net effect: the nested-type hierarchy this feature walks is guaranteed acyclic and finite by an existing, unrelated check. No extra depth limit or cycle guard needed.
Composite-agg-inside-nested-agg
Aggregating a nested field always requires wrapping it in an ES
nestedaggregation. Putting acompositeaggregation as its child is standard, well-supported Elasticsearch:{ "aggs": { "seasons": { "nested": { "path": "seasons_nested" }, "aggs": { "by_year": { "composite": { "size": 10, "sources": [ { "year": { "terms": { "field": "seasons_nested.year" } } } ] } } } } } }This is why the "no grouping above the target level" constraint matters:
after_keypagination only makes sense whencompositeis the single, top-most grouping construct in that branch. Aterms/compositeagg above it (e.g. bucketing by Team first) would produce oneafter_keyper parent bucket, making "page 2 of this nested list" ill-defined — which is the exact situation today'ssub_aggregationsare stuck in.Unverified: how composite grouping behaves on plain (non-nested) array fields within the target nested type — composite sources generally expect one value per document, and nested docs satisfy that, but an array-valued field inside the nested type might not.
Key code references
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/nested_sub_aggregation.rb— currentnested sub-aggregation query builder (terms + size cap, no composite).
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/composite_grouping_adapter.rb—existing composite/
after_keypagination (top-level only).elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/query.rb(filter_detail) — filtercomposition across parent/nested levels; the pattern ancestor filtering would reuse.
elasticgraph-schema_definition/lib/elastic_graph/schema_definition/mixins/supports_filtering_and_aggregation.rb— where sub-aggregation fields are defined without pagination support.
elasticgraph-graphql/spec/acceptance/sub_aggregations_spec.rb— current sub-aggregationbehavior and test coverage.
config/schema/teams.rb— schema used throughout this exploration (nested vs._objectfieldpairs,
nested_fields/nested_fields2duplicate-path case).All reactions