Skip to content

Standalone export of the advanced sklearn trainers emits a duplicate keyword argument instead of refusing duplicate parameter rows #8757

Description

@Yicong-Huang

What happened?

Exporting a workflow containing one of the four advanced sklearn trainers (KNN classifier/regressor, SVC, SVR) via POST /workflow-to-python produces a Python script that cannot be parsed when the operator's parameter table (paraList) contains two rows naming the same parameter. The emitted estimator call repeats the keyword, e.g. SVC(C = float ('1.0'),C = float ('2.0'),), and CPython raises SyntaxError: keyword argument repeated when the script is run. The export itself reports success (WorkflowToPythonSuccess); the failure only appears when the user runs the script.

What I expected: the export path refuses the duplicate the way the native path does. SklearnAdvancedBaseDesc.getParameter (used by the in-engine generatePythonCode) starts with requireDistinctParameters(paraList), whose own comment states the reason: "two rows naming one parameter emit that keyword twice and Python rejects the operator with a repeated-keyword SyntaxError. Refused here instead, where the workflow fails to compile with the parameter named." generateStandaloneCode() calls getParameterStandalone(paraList) directly and never invokes that guard, so the new export path lacks the backstop the PR itself added to the native path.

How to reproduce?

  1. Build a workflow JSON whose advanced trainer operator (e.g. SklearnAdvancedSVCTrainerOpDesc) has two paraList rows with the same parameter (e.g. two rows naming C). Note the form-level uniqueAmongRows validator blocks this interactively, so construct the plan by hand or via the API; stored workflows are not schema-validated on load, and the export endpoint takes the LogicalPlanPojo with no operator-property validation.
  2. POST the plan to /workflow-to-python.
  3. Run the returned script with Python: it fails with SyntaxError: keyword argument repeated.

For comparison, running the same operator natively fails at compile time with the parameter named (the require throws inside generatePythonCode, and getPhysicalOp embeds it as #EXCEPTION DURING CODE GENERATION:).

Version/Branch

1.4.0-incubating-SNAPSHOT (main)

Commit Hash (Optional)

Verified still present at acc6dd1139d936d849e2aaa6abb288f8ef3e98ad (main tip, 2026-09-29): requireDistinctParameters is defined at SklearnAdvancedBaseDesc.scala:266 and called only from getParameter (:286); generateStandaloneCode (:447-450) still bypasses it. Introduced by #8348 (merge 5127b78bc5dfd48a904894792c299ec72ba07ab9).

What browsers are you seeing the problem on?

n/a (backend export endpoint)

Relevant log output

n/a — the exporter returns success; the SyntaxError surfaces when running the emitted script.

Suggested fix

Call requireDistinctParameters(paraList) at the top of generateStandaloneCode() (or inside getParameterStandalone), and add a standalone regression test mirroring SklearnAdvancedBaseDescSpec.scala:359 ("refuse two rows setting one parameter"), which currently exercises only the native getParameter.

Activity

  1. added theissue type on Sep 29, 2026
  2. Yicong-Huang commented on Sep 29, 2026

    @Yicong-Huang
    ContributorAuthor

    @kz930 @aglinxinyuan — you two know this area best; would you take a look? Ignore if it is outside what you are working on.

  3. kz930 commented on Sep 29, 2026

    @kz930
    Contributor

    Confirmed on main: generateStandaloneCode never calls requireDistinctParameters. I'll take this and open a PR with the guard and a standalone regression test.

  4. added a commit that references this issue on Oct 3, 2026
    e0fe77a
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions