[SPARK-59684][SQL] pivot() on a struct column fails unless the pivot values are given explicitly - #58937
Draft
jiwen624 wants to merge 1 commit into
Draft
[SPARK-59684][SQL] pivot() on a struct column fails unless the pivot values are given explicitly#58937jiwen624 wants to merge 1 commit into
jiwen624 wants to merge 1 commit into
Conversation
jiwen624
force-pushed
the
SPARK-59684
branch
from
September 22, 2026 22:49
6c2e865 to
5396d7b
Compare
…lues are collected
jiwen624
force-pushed
the
SPARK-59684
branch
from
September 23, 2026 04:38
5396d7b to
c03dfad
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
RelationalGroupedDataset.collectPivotValues now returns typed Literals. It reads the pivot column's data type and converts each collected value with CatalystTypeConverters.createToCatalystConverter(dataType). Before, it returned raw external values and relied on Literal.apply to guess each type.
Both callers use the typed literals directly: classic pivot(Column) and the Spark Connect planner.
Why are the changes needed?
When pivot() is called without values, it collects the distinct values and converts them back into literals. Literal.apply can't do that for a Row, so pivoting by a struct column failed unless the values were given explicitly:
Before:
After:
The same thing broke array pivot columns and structs with UDT fields. Spark Connect failed too, with a raw UNSUPPORTED_FEATURE.LITERAL_TYPE error. The column's type is known when the values are collected, so this PR does the conversion there rather than widening Literal.apply.
Does this PR introduce any user-facing change?
Yes. pivot() without explicit values now works on struct columns, array columns, and structs with UDT fields, in both classic and Spark Connect. Other behavior is unchanged, including lit().
How was this patch tested?
Added UT cases.
Was this patch authored or co-authored using generative AI tooling?
Yes.