Skip to content

Core: Simplify logic in V4ManifestReader, update tests - #141

Closed
rdblue wants to merge 17 commits into
nastra:read-content-stats-from-manifestfrom
rdblue:read-content-stats-from-manifest
Closed

rdblue wants to merge 17 commits into
nastra:read-content-stats-from-manifestfrom
rdblue:read-content-stats-from-manifest

Conversation

@rdblue

@rdblue rdblue commented Sep 3, 2026 •

Copy link
Copy Markdown

This simplifies the logic in V4ManifestReader and updates to address the use cases from this comment.

This also rewrites the tests and makes some behavior changes that were found by updating the tests.

  • Tests are now all in the TestV4ManifestReader suite, which now also tests content stats cases. TestV4ManifestReaderStats is removed and deduplicated with the original reader suite.
  • Factory methods are renamed so that it is clear what is being produced; several cases where unpartitioned files were read as though partitioned are now fixed
  • Test cases are now reorganized so that each starts with a category: read, inheritance (mostly empty), projection, statsFilter (mostly empty), partitionFilter, validation, and resolution (path resolution)
  • Projection cases are tested directly by comparing the projection schema
  • Copying requested stats are tested separately to minimize cases that overlap with projection

Behavior changes:

  • metricsConfig is now required. Inferring the metrics config is dangerous and can easily drop metadata
  • tableLocation is used to configure the builder; NullPointerException is thrown if it is required but missing
  • Calling project(null) no longer resets the schema; this is invalid
  • Only requested stats should be copied

@rdblue rdblue changed the title Simplify projection logic using RestoreColumns, rewrite tests. Core: Simplify logic in V4ManifestReader, update tests Sep 3, 2026
// resolves stored locations against the table location
private TrackedFile copyResolved(TrackedFile trackedFile) {
TrackedFileStruct copy = (TrackedFileStruct) trackedFile.copy();
TrackedFileStruct copy = (TrackedFileStruct) trackedFile.copyWithStats(requestedStatsFieldIds);

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only requested stats should be copied.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reader also auto select stats for filter referenced fields. won't this change skip the filter stats during copy.

copyResolved is cloning the trackedFile after projection read from Parquet. do we need the redundant stats projection in copyResolved?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Empty requestedStatsFieldIds makes this "copy without stats". The default (no select/project/forScanPlanning) path still projects statsWriteSchema, so full stats are decoded and then dropped. Rewrite callers that used to keep all stats will lose them unless they pass projectStats for every field.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need the redundant stats projection in copyResolved?

Yes, this should return only the stats that are requested in the forScanPlanning case. The behavior in the existing scan planning path is to remove stats once they are used for filtering, unless they were requested by the caller.

Rewrite callers that used to keep all stats will lose them unless they pass projectStats for every field.

This is a good catch. When this is using the table's manifest schema with full content stats, all stats should be passed back to the caller, unless projectStats is called to narrow them.

Here's what I propose for each case:

  • forScanPlanning: project stats for all filter fields and projectStats fields, copy only the projectStats fields
  • No projection (rewrite case): project all stats in the table's manifest schema. If projectStats is called, pass the IDs to copyWithStats
  • select and project: all columns, including stats are projected based on the table's manifest schema. projectStats is not allowed because stats projection is controlled by the caller.

// content_stats is missing from the read schema when no stats are read
Types.NestedField statsField = readSchema.findField(TrackedFile.CONTENT_STATS_ID);
if (statsField != null) {
if (statsField != null && statsField.type().isStructType()) {

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rather than rewriting the schema to partition and stats fields when their type is unknown, this checks that the type is a struct before registering custom types for the fields. This is simpler overall and removes a schema rewrite.

this.tableLocation = tableLocation;
}

static Builder builder(

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Public static methods should be at the top of the file, not in the main body of the class.


/** Reader that reads a v4+ manifest file as {@link TrackedFile}s. */
class V4ManifestReader extends CloseableGroup implements CloseableIterable<TrackedFile> {
private static final Set<Integer> REQUIRED_COLUMN_IDS =

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required columns are now tracked as a set of IDs, with documentation.

* @param newTableLocation active table location
* @return this for method chaining
*/
Builder tableLocation(String newTableLocation) {

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not required. If it is needed but not supplied, path resolution throws a NullPointerException.

/** Sets the exact schema to read; used in place of {@link #select(Collection)}. */
/** Sets the exact schema to read; used in place of {@link #select(Iterable)}. */
Builder project(Schema newProjection) {
Preconditions.checkArgument(newProjection != null, "Invalid projection: null");

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no need to reset the projection by passing null here.


/**
* Sets the metrics config that determines which stats the manifest holds. Defaults to the
* config produced by the table's default metrics properties.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The metrics config is no longer defaulted. It looks like this default was intended to avoid rewriting test cases that did not pass the metrics config. This is a bad choice because metrics config is needed to determine the table's current manifest schema. Using an incorrect schema can easily lead to dropping content stats or storing extra content stats, and there is no way to detect the error.

scanMetrics);
}

private Map<Integer, Pair<Evaluator, StructProjection>> projectFilters() {

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Refactored to a private method.

private Schema readSchema(boolean includePartition) {
if (scanPlanning) {
// scan planning does not read the change-tracking fields omitted by SCAN_TYPE
Types.StructType statsProjection = StatsUtil.statsReadSchema(tableSchema, statsFieldIds());

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

statsFieldIds() combines requested stats IDs and filter IDs for the scan planning case.

if (scanPlanning || fieldIdsWithRequestedStats != null) {
return requiredStatsType;
}
Types.StructType tableStatsSchema = StatsUtil.statsWriteSchema(tableSchema, metricsConfig);

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All cases other than scan planning are projections of the table's current manifest schema, which is derived from the table schema and metrics config.

if (requestedProjection != null) {
return union(requiredStatsType, projectedStatsType(requestedProjection));
return RestoreColumns.restore(
tableManifestSchema, requestedProjection, idsToRestore(includePartition));

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RestoreColumns is a better way to ensure columns from the table manifest schema are present, instead of gathering all leaf field IDs and depending on select. The required IDs are known ahead of time and added to the projection.

import org.mockito.Mockito;

class TestV4ManifestReader {
private static final InputFile UNUSED_IN_FILE = Mockito.mock(InputFile.class);

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Used for cases where the file is not read.

TrackedFile.TRACKING.fieldId(),
Types.StructType.of(Tracking.STATUS, Tracking.SNAPSHOT_ID)))
.asStruct();
private static final Comparator<TrackedFile> FILE_COMPARATOR = comparator(FILE_VALIDATION_TYPE);

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This comparator is used to validate TrackedFile contents, excluding Tracking because tracking has inherited fields that do not match the file that was written.

"s3://bucket/table/id=2/eq-deletes-b.parquet", idPartition(2));
private static final TrackedFile DATA_MANIFEST_REF =
manifestRef(FileContent.DATA_MANIFEST, "data-leaf.parquet");
manifestRef(FileContent.DATA_MANIFEST, "s3://bucket/table/data-leaf.parquet");

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All tests that are not related to path resolution use full URIs. Path resolution tests embed locations so that the test cases are more readable.

@ParameterizedTest
@FieldSource("MANIFEST_FORMATS")
public void equalityDeleteRoundTrip(FileFormat format) throws IOException {
public void readEqualityDelete(FileFormat format) throws IOException {

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There are now read tests for data file, delete file, and manifest file with all expected fields set. There are also read tests to validate that requested stats are copied correctly.

fileWithStatus(EntryStatus.DELETED, "s3://bucket/deleted.parquet"),
fileWithStatus(EntryStatus.REPLACED, "s3://bucket/replaced.parquet"));
@Test
public void projectionFullByDefault() {

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Most projection tests get the reader's projection and validate it directly. This was important with stats because stats may be removed at the copy phase.

}
Types.StructType readSchema = builder.build().readSchema().asStruct();

assertThat(readSchema.field(TrackedFile.LOCATION.fieldId()))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests for select and project don't validate all fields. They validate that the projected field is present, at least one required field is present, and that fields needed for filters are (or are not) present.

import org.apache.iceberg.relocated.com.google.common.collect.Lists;
import org.apache.iceberg.schema.SchemaWithPartnerVisitor;

public class RestoreColumns extends SchemaWithPartnerVisitor<Type, Type> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should this be package private?

@rdblue rdblue Sep 10, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It could be, but this is in the types package and we need to use it in the reader, which isn't in the same package. We have two options here: add a method to TypeUtil or expose RestoreColumns.restore.

We have precedent for both of these patterns in the codebase (like Binder.bind) and I've opted to use the second one because TypeUtil is getting to be large.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was initially thinking the same thing but SchemaWithPartnerVisitor itself is public, so by making RestoreColumns public we don't accidentally risk exposing methods that shouldn't be public. Therefore I think it's fine to keep RestoreColumns as public

Comment thread api/src/main/java/org/apache/iceberg/expressions/Binder.java
// resolves stored locations against the table location
private TrackedFile copyResolved(TrackedFile trackedFile) {
TrackedFileStruct copy = (TrackedFileStruct) trackedFile.copy();
TrackedFileStruct copy = (TrackedFileStruct) trackedFile.copyWithStats(requestedStatsFieldIds);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reader also auto select stats for filter referenced fields. won't this change skip the filter stats during copy.

copyResolved is cloning the trackedFile after projection read from Parquet. do we need the redundant stats projection in copyResolved?

// resolves stored locations against the table location
private TrackedFile copyResolved(TrackedFile trackedFile) {
TrackedFileStruct copy = (TrackedFileStruct) trackedFile.copy();
TrackedFileStruct copy = (TrackedFileStruct) trackedFile.copyWithStats(requestedStatsFieldIds);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Empty requestedStatsFieldIds makes this "copy without stats". The default (no select/project/forScanPlanning) path still projects statsWriteSchema, so full stats are decoded and then dropped. Rewrite callers that used to keep all stats will lose them unless they pass projectStats for every field.

}

@Test
void unknownToStructProjection() {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need the test coverage for the other direction: structToUnknownProjection()?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think so. Structs can't change type to unknown, but unknown can be evolved to a struct.

I'm also not sure that we would be using a schema that was evolved, rather than just restoring columns in a projection. I'd leave this as-is for now.

*/
Builder tableLocation(String newTableLocation) {
Preconditions.checkArgument(newTableLocation != null, "Invalid table location: null");
this.tableLocation = newTableLocation;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The previous builder stored LocationUtil.stripTrailingSlash(tableLocation) at line 220. resolveLocation documents that tableLocation must not end with /, or the result gets //. Worth stripping here again.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was fixed in a concurrent update. It is not the responsibility of the manifest reader to alter the table location.


@ParameterizedTest
@FieldSource("MANIFEST_FORMATS")
void geoStatsAreCorrectWithContainerReuse(FileFormat format) throws IOException {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tests around variant/geo with/without container reuse reproduced actual bugs but appear to not be copied over to TestV4ManifestReader and seem to be removed

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason for removing these is that it isn't the right place to test the fully copy behavior of TrackedFile. We have tests in TrackedFile, ContentStats, and FieldStats to validate the copy behavior for each type.

In this suite, what we need to test is that copy or copyWithStats is called.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fair point. let me extract those tests into a separate PR and test the container reuse bugs in the files you mentioned so that we have proper coverage for those scenarios

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@rdblue I've pulled the fixes and tests out into apache#18060. Please take a look

import org.junit.jupiter.params.ParameterizedTest;
import org.junit.jupiter.params.provider.FieldSource;

class TestV4ManifestReaderStats {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the reason I moved the stats-related tests into TestV4ManifestReaderStats was because TestV4ManifestReader is already really large and to also align with how we tested these for v1-v3, where we also have TestManifestReader/TestManifestReaderStats

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that the separation left bugs in the reader suite. And there were a ton of tests here that weren't needed because they were testing other components, like the stats tests for specific types.


@ParameterizedTest
@FieldSource("MANIFEST_FORMATS")
void readStatsForNestedFields(FileFormat format) throws IOException {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks like this hasn't been moved over

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Field stats are all tracked at the same level. Does it matter whether the original field was nested or not? I don't see why this would be any different.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the main motivation was to be explicit and verify that this behavior actually works instead of assuming that it would work. We can also move this to a different place/class if you think it's at the wrong place here


@ParameterizedTest
@FieldSource("MANIFEST_FORMATS")
void selectStatsByName(FileFormat format) throws IOException {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would be good to keep this test as well

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't stats columns fall in the same category as any other requested column? Why would this be different?

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

they do but I feel like we'd rather have an explicit test to make sure this does what we think it does rather than assuming it would do the right thing

@rdblue

rdblue commented Sep 10, 2026

Copy link
Copy Markdown
Author

@nastra and @stevenzwu, I've updated this. The main change was to only filter stats when forScanPlanning is called, or if projectStats is called when projecting the whole stats structure (to narrow stats). The select and project cases both do not allow projectStats.

The reason this took longer than expected was that some tests started failing when the stats were empty. I tracked down the problem and it was that the schema used to build the StructComparator didn't match the schema used for reading. I ended up needing to produce different comparators that check the API rather than comparing structs by position. I think this is now a better way to validate, but it took a bit more time to track it down.

Comment thread core/src/main/java/org/apache/iceberg/types/RestoreColumns.java

return tableLocation + PATH_SEPARATOR + location;
// call toString to cause a NullPointerException if the string is null
return tableLocation.toString() + PATH_SEPARATOR + location;

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to use Preconditions.checkArgument to be more explicit here or did you want to avoid using it because this method is on the hot path?

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's also a weird idiom in a public util method that we don't really do anywhere in the codebase. The other issue with this is that not every JDK version would actually produce a helpful NPE with an actual message (usually < JDK 17). But this can be even true for JDK >= 17 depending on the JIT compiler. Also the message can be different depending on the actual JDK vendor being used. That being said, I think we're relying here too much on the underlying behavior of a particular JDK

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the behavior that we are relying on in a particular JDK? Won't any JDK need to throw a NullPointerException if tableLocation is null?

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the thing that's relying on the JDK is the actual error msg and its format. In certain cases the compiler can just omit the underlying error msg. Other JDK implementations can choose to have a different msg format or also completely omit it

private Set<String> requestedColumns = null;
private Schema requestedProjection = null;
private Set<Integer> fieldIdsWithRequestedStats = null;
private Set<Integer> requestedStatsFieldIds = null;

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: I still feel like fieldIdsWithRequestedStats slightly better describes that these are field IDs for which stats where requested. requestedStatsFieldIds implies that those are the actual stats field IDs, but the user of the projectStats API always supplies actual table field IDs and so to me reading requestedStatsFieldIds always requires one additional mental step in order to be sure this actually carries table field IDs. Would be good to know how others read this field name

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that this distinction is going to be lost for anyone reading this, so we should use the name that is more straightforward. We are keeping the stats field IDs encapsulated in stats utils as much as possible to avoid this kind of issue.

STATS_TYPE.fieldType("data").asStructType(), "a", "z", false, RECORD_COUNT, 20, 0, null);
private static final ContentStatsStruct CONTENT_STATS = new ContentStatsStruct(STATS_TYPE);

static {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's a bit weird to have a static initializer in tests. I think it's probably better to have a method annotated with @BeforeAll

@rdblue rdblue closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants