Skip to content

[SPARK-59684][SQL] pivot() on a struct column fails unless the pivot values are given explicitly - #58937

Draft
jiwen624 wants to merge 1 commit into
apache:masterfrom
jiwen624:SPARK-59684
Draft

jiwen624 wants to merge 1 commit into
apache:masterfrom
jiwen624:SPARK-59684

Conversation

@jiwen624

@jiwen624 jiwen624 commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this pull request?

RelationalGroupedDataset.collectPivotValues now returns typed Literals. It reads the pivot column's data type and converts each collected value with CatalystTypeConverters.createToCatalystConverter(dataType). Before, it returned raw external values and relied on Literal.apply to guess each type.

Both callers use the typed literals directly: classic pivot(Column) and the Spark Connect planner.

Why are the changes needed?

When pivot() is called without values, it collects the distinct values and converts them back into literals. Literal.apply can't do that for a Row, so pivoting by a struct column failed unless the values were given explicitly:

val df = Seq(1.0, 2.0).toDF("v").selectExpr("v", "struct(v, v) AS s")
df.groupBy("v").pivot("s").count().orderBy("v").show()

Before:

org.apache.spark.SparkRuntimeException: [UNSUPPORTED_FEATURE.PIVOT_TYPE] The feature is not supported: Pivoting by the value '[1.0,1.0]' of the column data type "STRUCT<v: DOUBLE NOT NULL, v: DOUBLE NOT NULL>". SQLSTATE: 0A000

After:

+---+----------+----------+
|  v|{1.0, 1.0}|{2.0, 2.0}|
+---+----------+----------+
|1.0|         1|      NULL|
|2.0|      NULL|         1|
+---+----------+----------+

The same thing broke array pivot columns and structs with UDT fields. Spark Connect failed too, with a raw UNSUPPORTED_FEATURE.LITERAL_TYPE error. The column's type is known when the values are collected, so this PR does the conversion there rather than widening Literal.apply.

Does this PR introduce any user-facing change?

Yes. pivot() without explicit values now works on struct columns, array columns, and structs with UDT fields, in both classic and Spark Connect. Other behavior is unchanged, including lit().

How was this patch tested?

Added UT cases.

Was this patch authored or co-authored using generative AI tooling?

Yes.

@jiwen624 jiwen624 changed the title [WIP][SPARK-59684][SQL] pivot() on a struct column fails unless the pivot values are given explicitly [SPARK-59684][SQL] pivot() on a struct column fails unless the pivot values are given explicitly Sep 22, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant