[BEAM-3227] Consider sharing Udf/SdkFunctionSpec records via pointer - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Sub-task
Status: Resolved
Priority: P2
Resolution: Won't Fix
Affects Version/s: None
Fix Version/s: Not applicable
Component/s: beam-model
Labels:
None

Description

Coders are stored by pointer, because they are often repeated and a common source of huge pipeline descriptions.

We considered doing the same for all UDFs but decided not to, based on the logic that they are not as often identical and will rarely implement the equals() needed to actually share encoded versions.

However, in the presence of generated code, it is very likely that DoFns and CombineFns are repeated, and also much more likely that they have meaningful equals(), so there could be size savings.

None of this is terribly important for storage or transmission, but has more to do with arbitrary and small size limits that occur in some API frameworks or database column types.

Attachments

Activity

People

Assignee:: Unassigned

Reporter:: Kenneth Knowles

Votes:: 0 Vote for this issue

Watchers:: 2 Start watching this issue

Dates

Created:: 20/Nov/17 16:22

Updated:: 16/May/20 14:16

Resolved:: 23/Apr/20 16:39