Problem
sub_aggregations let a client aggregate nested collections (e.g. Team.seasons_nested), but only as a child of a parent-level aggregation (team_aggregations { nodes { sub_aggregations { seasons_nested { ... } } } }).
No real pagination. Sub-aggregation connections expose first and has_next_page but no after/cursors. They're built on a plain terms aggregation with a size cap, not a composite aggregation. Composite's cursor-based pagination only works today for top-level aggregations, where there's a single grouping level with nothing above it. Once a nested aggregation sits under a parent grouping, a cursor per parent bucket isn't well-defined, so cursor pagination isn't offered at all. This is the core problem this design addresses.
Proposed feature
Offer nested types as root-level aggregation query fields:
query {
teamSeasonPlayerAggregations {
pageInfo { hasNextPage endCursor }
nodes {
groupedBy { ... }
countDetail { ... }
aggregatedValues { ... }
}
}
}
The query schema would guarantee no grouping happens above the target nested level for this field because it aggregates directly over the nested collection with no parent-level bucketing in the way. That's what makes it valid to use a real composite aggregation (nested inside the nested wrapper), giving after-based cursor pagination consistent with top-level aggregation.
This is additive in that existing sub_aggregations are untouched and keep today's first-only behavior. Root-level fields are a new, separate entry point into the same nested collections.
As a secondary benefit, these new root-level fields also give clients a shorter path to a deeply-nested collection. To aggregate Player rows nested three levels down (Team → TeamSeason → Player), a client using today's sub_aggregations must traverse an extra nodes { sub_aggregations { ... } } } layer per intermediate level, even when they have no interest in grouping at Team or Season at all (an intermediate level can be left ungrouped, but its nodes/sub_aggregations wrapper still has to be written out); a root field skips straight to the target level. This is a convenience; the query already works today, just more verbosely.
Ancestor filtering
For a root field aggregating Player three levels deep — Team --(seasons_nested)--> TeamSeason --(some_middle_nested)--> Middle --(players_nested)--> Player, (Middle doesn't exist in the teams schema today but illustrates a deep nesting case) clients may want to filter on ancestor fields (Team, TeamSeason, Middle), not just Player fields.
Proposal: the field's own type (Player) filters are specified inline in filter, exactly as it would for a sub_aggregations field's filter today. Every ancestor, including the root, gets its own key, one per step in the path, nested to mirror the actual path structure:
filter: {
team: {
league: { equal_to_any_of: ["NFL"] }
seasons_nested: {
year: { gt: 2020 }
some_middle_nested: {
someMiddleField: { ... }
}
}
}
name: { equal_to_any_of: ["Smith"] } # Player, inline (target type)
}
team is the outermost/topmost ancestor key (the root document type); seasons_nested is the field Team uses to reach TeamSeason, and some_middle_nested is the field TeamSeason uses to reach Middle — each nests inside the previous ancestor's block since each ancestor is only reachable through the one before it.
The path from root to target nests one level at a time, each key being the field-path segment used to reach that level; only the target type's own fields are inline. This avoids two problems with a flatter alternative: field name collisions between the root type and the target type (both being flattened into the same object could produce duplicate GraphQL field names), and inconsistent treatment of the root ancestor vs. every other ancestor.
Keying ancestor blocks by field-path segment (rather than type name) keeps them unambiguous even when two ancestor levels share a type, or the same type is reachable via two differently-named paths. The teams schema's nested_fields/nested_fields2 is exactly this case — each duplicate path already has a distinct field name to key off of. These key names should be configurable via the schema DSL, consistent with how other schema element names are already customizable.
Naming / schema surface
Every distinct nested path to a given type needs its own uniquely-named root field. A naming scheme derived automatically from type names alone has potential problems: the teams schema's nested_fields/nested_fields2 (two distinct paths to the same TeamNestedFields type) would collide if the root field name were built purely from the target type.
Proposal: root-level aggregation fields are specified in the schema. A schema author explicitly opts-in by declaring which nested paths get promoted to a root field, and supplies the root field name in doing so, rather than any name being auto-derived from the path or type name. Since names are author-supplied, schema-build-time validation should reject duplicate root field names across different declarations, so a naming mistake surfaces immediately as a build error.
Proposed DSL shape — a declaration on the innermost nested field, naming the root field it should be promoted to:
t.field "players_nested", "[Player!]!" do |f|
f.mapping type: "nested"
f.root_aggregation_field "team_season_player_aggregations"
end
Known gap: no client-specified sorting (acknowledged, not addressed here)
Neither top-level aggregations nor sub-aggregations support client-specified sorting today — ordering is implicit, driven by grouping keys (ascending, via composite sources for top-level;
doc_count descending, then re-sorted client-side for stability, for sub-aggregations). This is a documented design choice (config/site/query-api/aggregations.md), not an oversight.
Why it matters here: pagination is often only useful paired with sorting — e.g. "top 5 seasons by win count" needs sorting by an aggregated_values metric, not by grouping key. Adding cursor
pagination to root-level fields doesn't make that possible on its own; it only lets you page through buckets in grouping-key order.
Sorting by a metric would need an ES bucket_sort pipeline aggregation, which interacts awkwardly with composite's after_key traversal (which assumes ascending key order) — combining the two isn't a drop-in change and would need its own design.
Decision: sorting is out of scope for this design. Root-level aggregation fields ship without client-specified sorting, matching current behavior everywhere else.
Non-goals / deferred
- Ancestor filtering may ship as a follow-on to a first cut that only filters within the target nested path and its descendants. The filter-key design above should anticipate ancestor filtering from the start so the later extension isn't a breaking change.
- Cross-sibling-nested filtering / reverse_nested-based root counts (e.g. "count distinct Teams that have a Player matching some condition") — out of scope here, potential future extension.
- Client-specified sorting — see above; out of scope here.
Open questions
- Whether composite grouping on non-nested array fields within the target nested type works correctly (unverified).
Appendix: other aspects considered
reverse_nested isn't needed (for the base feature)
It seemed like ancestor filtering (e.g. filtering a Player root field by a TeamSeason or Team field) would need Elasticsearch's reverse_nested aggregation, since the query "starts" at the nested level. It doesn't — reverse_nested isn't used anywhere in the codebase today, and it isn't required here either. The root field is a resolver-level convenience; the actual ES request still starts at the real index root and nests downward as usual:
filter(team) → nested(season) → filter(season) → nested(player) → composite(player)
This is the same filter-composition pattern Aggregation::Query#filter_detail already builds for today's parent-filtered sub-aggregations. The only new part is that the resolver skips exposing the intermediate Team/Season grouping levels and returns the innermost composite result directly.
reverse_nested would only be needed for a different class of query we're not targeting — e.g. "count distinct Teams that have a Player matching some condition," or filtering by a sibling nested collection at the same depth. This is not in scope.
Recursive/self-referential nested types: not a concern
A type like Category with a self-referential nested field (subcategories: [Category!]!, allowing arbitrary depth) would need special handling for this feature, but ElasticGraph already disallows that pattern at schema-build time regardless — it rejects any cycle among plain/nested/embedded fields (schema_definition/results.rb, check_for_circular_dependencies!). The only legal self-reference is through relates_to_one/relates_to_many relationship fields, which are a different mechanism and are already excluded from sub-aggregation path generation.
Net effect: the nested-type hierarchy this feature walks is guaranteed acyclic and finite by an existing, unrelated check. No extra depth limit or cycle guard needed.
Composite-agg-inside-nested-agg
Aggregating a nested field always requires wrapping it in an ES nested aggregation. Putting a composite aggregation as its child is standard, well-supported Elasticsearch:
{
"aggs": {
"seasons": {
"nested": { "path": "seasons_nested" },
"aggs": {
"by_year": {
"composite": {
"size": 10,
"sources": [
{ "year": { "terms": { "field": "seasons_nested.year" } } }
]
}
}
}
}
}
}
This is why the "no grouping above the target level" constraint matters: after_key pagination only makes sense when composite is the single, top-most grouping construct in that branch. A terms/composite agg above it (e.g. bucketing by Team first) would produce one after_key per parent bucket, making "page 2 of this nested list" ill-defined — which is the exact situation today's sub_aggregations are stuck in.
Unverified: how composite grouping behaves on plain (non-nested) array fields within the target nested type — composite sources generally expect one value per document, and nested docs satisfy that, but an array-valued field inside the nested type might not.
Key code references
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/nested_sub_aggregation.rb — current
nested sub-aggregation query builder (terms + size cap, no composite).
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/composite_grouping_adapter.rb —
existing composite/after_key pagination (top-level only).
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/query.rb (filter_detail) — filter
composition across parent/nested levels; the pattern ancestor filtering would reuse.
elasticgraph-schema_definition/lib/elastic_graph/schema_definition/mixins/supports_filtering_and_aggregation.rb
— where sub-aggregation fields are defined without pagination support.
elasticgraph-graphql/spec/acceptance/sub_aggregations_spec.rb — current sub-aggregation
behavior and test coverage.
config/schema/teams.rb — schema used throughout this exploration (nested vs. _object field
pairs, nested_fields/nested_fields2 duplicate-path case).
Problem
sub_aggregationslet a client aggregate nested collections (e.g.Team.seasons_nested), but only as a child of a parent-level aggregation (team_aggregations { nodes { sub_aggregations { seasons_nested { ... } } } }).No real pagination. Sub-aggregation connections expose
firstandhas_next_pagebut noafter/cursors. They're built on a plain terms aggregation with asizecap, not acompositeaggregation. Composite's cursor-based pagination only works today for top-level aggregations, where there's a single grouping level with nothing above it. Once a nested aggregation sits under a parent grouping, a cursor per parent bucket isn't well-defined, so cursor pagination isn't offered at all. This is the core problem this design addresses.Proposed feature
Offer nested types as root-level aggregation query fields:
The query schema would guarantee no grouping happens above the target nested level for this field because it aggregates directly over the nested collection with no parent-level bucketing in the way. That's what makes it valid to use a real
compositeaggregation (nested inside thenestedwrapper), givingafter-based cursor pagination consistent with top-level aggregation.This is additive in that existing
sub_aggregationsare untouched and keep today'sfirst-only behavior. Root-level fields are a new, separate entry point into the same nested collections.As a secondary benefit, these new root-level fields also give clients a shorter path to a deeply-nested collection. To aggregate
Playerrows nested three levels down (Team→TeamSeason→Player), a client using today'ssub_aggregationsmust traverse an extranodes { sub_aggregations { ... } } }layer per intermediate level, even when they have no interest in grouping at Team or Season at all (an intermediate level can be left ungrouped, but itsnodes/sub_aggregationswrapper still has to be written out); a root field skips straight to the target level. This is a convenience; the query already works today, just more verbosely.Ancestor filtering
For a root field aggregating
Playerthree levels deep —Team --(seasons_nested)--> TeamSeason --(some_middle_nested)--> Middle --(players_nested)--> Player, (Middledoesn't exist in theteamsschema today but illustrates a deep nesting case) clients may want to filter on ancestor fields (Team, TeamSeason, Middle), not just Player fields.Proposal: the field's own type (
Player) filters are specified inline infilter, exactly as it would for asub_aggregationsfield's filter today. Every ancestor, including the root, gets its own key, one per step in the path, nested to mirror the actual path structure:teamis the outermost/topmost ancestor key (the root document type);seasons_nestedis the fieldTeamuses to reachTeamSeason, andsome_middle_nestedis the fieldTeamSeasonuses to reachMiddle— each nests inside the previous ancestor's block since each ancestor is only reachable through the one before it.The path from root to target nests one level at a time, each key being the field-path segment used to reach that level; only the target type's own fields are inline. This avoids two problems with a flatter alternative: field name collisions between the root type and the target type (both being flattened into the same object could produce duplicate GraphQL field names), and inconsistent treatment of the root ancestor vs. every other ancestor.
Keying ancestor blocks by field-path segment (rather than type name) keeps them unambiguous even when two ancestor levels share a type, or the same type is reachable via two differently-named paths. The
teamsschema'snested_fields/nested_fields2is exactly this case — each duplicate path already has a distinct field name to key off of. These key names should be configurable via the schema DSL, consistent with how other schema element names are already customizable.Naming / schema surface
Every distinct nested path to a given type needs its own uniquely-named root field. A naming scheme derived automatically from type names alone has potential problems: the
teamsschema'snested_fields/nested_fields2(two distinct paths to the sameTeamNestedFieldstype) would collide if the root field name were built purely from the target type.Proposal: root-level aggregation fields are specified in the schema. A schema author explicitly opts-in by declaring which nested paths get promoted to a root field, and supplies the root field name in doing so, rather than any name being auto-derived from the path or type name. Since names are author-supplied, schema-build-time validation should reject duplicate root field names across different declarations, so a naming mistake surfaces immediately as a build error.
Proposed DSL shape — a declaration on the innermost nested field, naming the root field it should be promoted to:
Known gap: no client-specified sorting (acknowledged, not addressed here)
Neither top-level aggregations nor sub-aggregations support client-specified sorting today — ordering is implicit, driven by grouping keys (ascending, via composite
sourcesfor top-level;doc_countdescending, then re-sorted client-side for stability, for sub-aggregations). This is a documented design choice (config/site/query-api/aggregations.md), not an oversight.Why it matters here: pagination is often only useful paired with sorting — e.g. "top 5 seasons by win count" needs sorting by an
aggregated_valuesmetric, not by grouping key. Adding cursorpagination to root-level fields doesn't make that possible on its own; it only lets you page through buckets in grouping-key order.
Sorting by a metric would need an ES
bucket_sortpipeline aggregation, which interacts awkwardly with composite'safter_keytraversal (which assumes ascending key order) — combining the two isn't a drop-in change and would need its own design.Decision: sorting is out of scope for this design. Root-level aggregation fields ship without client-specified sorting, matching current behavior everywhere else.
Non-goals / deferred
Open questions
Appendix: other aspects considered
reverse_nestedisn't needed (for the base feature)It seemed like ancestor filtering (e.g. filtering a
Playerroot field by aTeamSeasonorTeamfield) would need Elasticsearch'sreverse_nestedaggregation, since the query "starts" at the nested level. It doesn't —reverse_nestedisn't used anywhere in the codebase today, and it isn't required here either. The root field is a resolver-level convenience; the actual ES request still starts at the real index root and nests downward as usual:This is the same filter-composition pattern
Aggregation::Query#filter_detailalready builds for today's parent-filtered sub-aggregations. The only new part is that the resolver skips exposing the intermediate Team/Season grouping levels and returns the innermost composite result directly.reverse_nestedwould only be needed for a different class of query we're not targeting — e.g. "count distinct Teams that have a Player matching some condition," or filtering by a sibling nested collection at the same depth. This is not in scope.Recursive/self-referential nested types: not a concern
A type like
Categorywith a self-referential nested field (subcategories: [Category!]!, allowing arbitrary depth) would need special handling for this feature, but ElasticGraph already disallows that pattern at schema-build time regardless — it rejects any cycle among plain/nested/embedded fields (schema_definition/results.rb,check_for_circular_dependencies!). The only legal self-reference is throughrelates_to_one/relates_to_manyrelationship fields, which are a different mechanism and are already excluded from sub-aggregation path generation.Net effect: the nested-type hierarchy this feature walks is guaranteed acyclic and finite by an existing, unrelated check. No extra depth limit or cycle guard needed.
Composite-agg-inside-nested-agg
Aggregating a nested field always requires wrapping it in an ES
nestedaggregation. Putting acompositeaggregation as its child is standard, well-supported Elasticsearch:{ "aggs": { "seasons": { "nested": { "path": "seasons_nested" }, "aggs": { "by_year": { "composite": { "size": 10, "sources": [ { "year": { "terms": { "field": "seasons_nested.year" } } } ] } } } } } }This is why the "no grouping above the target level" constraint matters:
after_keypagination only makes sense whencompositeis the single, top-most grouping construct in that branch. Aterms/compositeagg above it (e.g. bucketing by Team first) would produce oneafter_keyper parent bucket, making "page 2 of this nested list" ill-defined — which is the exact situation today'ssub_aggregationsare stuck in.Unverified: how composite grouping behaves on plain (non-nested) array fields within the target nested type — composite sources generally expect one value per document, and nested docs satisfy that, but an array-valued field inside the nested type might not.
Key code references
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/nested_sub_aggregation.rb— currentnested sub-aggregation query builder (terms + size cap, no composite).
elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/composite_grouping_adapter.rb—existing composite/
after_keypagination (top-level only).elasticgraph-graphql/lib/elastic_graph/graphql/aggregation/query.rb(filter_detail) — filtercomposition across parent/nested levels; the pattern ancestor filtering would reuse.
elasticgraph-schema_definition/lib/elastic_graph/schema_definition/mixins/supports_filtering_and_aggregation.rb— where sub-aggregation fields are defined without pagination support.
elasticgraph-graphql/spec/acceptance/sub_aggregations_spec.rb— current sub-aggregationbehavior and test coverage.
config/schema/teams.rb— schema used throughout this exploration (nested vs._objectfieldpairs,
nested_fields/nested_fields2duplicate-path case).