# Databricks hires FDEs on Spark, but Unity Catalog and Genie are left off the requirements list

> The job posting asks for Spark, but the product documentation shows that Genie only gives correct answers when access rights, row filters and business rules are done carefully.

Original: https://fdetimes.net/en/tools/databricks-fde-spark-unity-catalog-genie/

The Databricks job posting for a Forward Deployed Engineer in Seoul states plainly that candidates need deep experience in distributed computing with Apache Spark and an understanding of Spark's runtime internals. Genie and Unity Catalog, by contrast, appear only in the company introduction, not in the list of requirements.

Spark is what you need to prepare to get through the hiring process. According to the official documentation, though, deploying Unity Catalog and Genie is a matter of access rights, metadata and business rules.

## Spark: learn it well enough to explain a slow job

The Seoul posting asks for 7+ years in data engineering, data platforms and analytics, or software engineering; coding ability in Python, Scala or JavaScript/TypeScript; and roughly 20% of time spent travelling to customer sites. The Senior FDE posting for National Security keeps the same Spark requirement, lowers the bar to 6+ years and adds practical knowledge of MLOps, ML/AI models and AI APIs.

"Runtime internals" sounds vague, but it can be reduced to three concrete things. The first is the shuffle: joins and aggregations force data to move between nodes, and `spark.sql.shuffle.partitions` splits that work into 200 partitions by default.

The second is Adaptive Query Execution, enabled by default since Spark 3.2.0. AQE uses runtime statistics to coalesce post-shuffle partitions that are too small and to split skewed partitions in sort-merge joins. The third is the broadcast join: a table smaller than `spark.sql.autoBroadcastJoinThreshold`, 10 MB by default, is sent to every worker instead of being shuffled.

The way to practise is to make a job slow on purpose and then read the Spark UI. Join a large table with a small one, open the SQL tab and click Details to see whether the physical plan uses BroadcastHashJoin or SortMergeJoin. Disable broadcasting by setting the threshold to -1, run it again and compare the timings.

Then open the detail page for the slowest stage. The Summary Metrics table shows the distribution of task run times: if the slowest task takes many times longer than the median, the data is probably skewed. The Shuffle Read Size and Shuffle spill (disk) columns show how much data moved and whether it spilled to disk.

Practise describing what you saw using the actual numbers in the UI. In an interview, that is the most concrete way to show you understand how Spark works underneath.

## Unity Catalog answers the question "who can see what?"

Databricks defines Unity Catalog as a unified governance layer for data and AI, built into the platform. Every object lives in a three-level namespace, `catalog.schema.object`. Tables and volumes can be managed, meaning Unity Catalog handles both governance and the lifecycle of the underlying storage files, or external.

According to Microsoft Learn, every Azure Databricks workspace created after 9 November 2023 has Unity Catalog enabled automatically, and Unity Catalog also has an open-source edition. So at a customer running a newer workspace, you will probably be working with it from day one.

The documentation divides Unity Catalog's capabilities into several areas, of which three are worth learning first: permissions, lineage and audit. Permissions are handled through privileges, attribute-based access control (ABAC) policies, row filters, column filters and workspace bindings.

Lineage records the path from source data to models, services and dashboards without manual configuration, while every data access and system activity is captured in a system table dedicated to audit logs, which is what the compliance team will need.

## Example: each branch sees only its own data

Imagine a retail chain with a table `sales.core.orders` that has a `region` column. The requirement: the analyst group in Vietnam may see only Vietnamese orders, while the admin group sees everything. The first step is to grant read access:

```sql
GRANT USE CATALOG ON CATALOG sales TO `analysts_vn`;
GRANT USE SCHEMA ON SCHEMA sales.core TO `analysts_vn`;
GRANT SELECT ON TABLE sales.core.orders TO `analysts_vn`;
```

With GRANT SELECT alone, this group can still read every row. To restrict by region, you write a function that returns true or false for each row, then attach that function to the table:

```sql
CREATE FUNCTION sales.gov.region_filter(region STRING)
RETURN IF(is_account_group_member('admins'), true, region = 'VN');

ALTER TABLE sales.core.orders
SET ROW FILTER sales.gov.region_filter ON (region);
```

From then on, the same `SELECT * FROM sales.core.orders` returns two different results depending on who runs it. The access logic lives on the table rather than being scattered across individual dashboards. At a customer, sit down with the data owner to agree the "who sees which rows" table first, and only then write the function.

Two details from the documentation are worth remembering. The function's parameter type must match the column type: if they differ and ANSI mode is off, values that cannot be cast become NULL, and the filter may return the entire table without raising an error.

Also, when the same rule is needed across many tables, Databricks recommends using an ABAC policy attached at the catalog or schema level rather than attaching row filters table by table.

## Genie sits on top of Unity Catalog, so wrong permissions mean wrong answers

According to documentation from September 2026, Genie is now a product family comprising Genie One, Genie Agents and Genie Code. Genie Agents are environments organised by business domain, where data teams configure trusted data, metrics and business rules that Genie One relies on to answer questions.

The naming has also changed: the current overview page no longer uses the term "Genie space". When following an older tutorial, check its update date first.

Databricks states that every Genie answer is grounded in the organisation's data and passes through Unity Catalog's governance layer.

The practical consequence: the row filter in the example above also determines what Genie users in Vietnam see. If the permissions are done carelessly, Genie will deliver that mistake straight to business users in the form of a highly convincing answer.

Building a Genie Agent is a configuration job: choose datasets, write sample queries, write instructions. Quality is then tuned through metrics, business rules and verified answers. For instance, if the customer defines "net revenue" as excluding returned goods, that must be a business rule, and the most frequently asked question of the week should have a verified answer.

According to the documentation, users can use Genie One and Genie Agents free of charge until 31 January 2027; service principals are not covered. That makes now a convenient time to get hands-on practice.

**Key point:** A data question-and-answer assistant is only trustworthy when the permission layer beneath it has been built carefully.

## What to learn first, and what to put on your CV

The sensible order is Spark and the Spark UI first, because that is the way in. Next comes Unity Catalog governance SQL: GRANT, row filters, column filters, reading lineage and audit logs. Genie Agents come last, because they depend on the other two layers.

On your CV, instead of writing "knows Databricks", describe specific work: "identified skew in a join using the Spark UI, cutting job time from X to Y", or "designed regional row filters for an orders table, verified through audit logs".

When reading a job description, do not skip the products that appear only in the company introduction. They are usually what you will deploy, and with Genie, the quality of the answers depends on the permission layer underneath.

**Try this week:**

- Join a large table with a small one, read the physical plan in the Spark UI's SQL tab, then set autoBroadcastJoinThreshold to -1 and compare run times.
- Create an orders table with a region column in a Databricks workspace, then write the GRANT statements, the row filter function and ALTER TABLE SET ROW FILTER as in the example above.
- Log in as two users from two different groups, run the same SELECT, and compare how many rows each receives.
- Open the audit log table in system tables, filter by the users and time window you just tested, and note which events were actually recorded.

## Sources

- [What is Unity Catalog? | Databricks on AWS](https://docs.databricks.com/aws/en/data-governance/unity-catalog/)

- [What is Unity Catalog? - Azure Databricks | Microsoft Learn](https://learn.microsoft.com/en-us/azure/databricks/data-governance/unity-catalog/)

- [Genie - Azure Databricks | Microsoft Learn](https://learn.microsoft.com/en-us/azure/databricks/genie/)

- [Genie | Databricks on AWS](https://docs.databricks.com/en/genie/index.html)

- [Forward Deployed Engineer - Databricks](https://www.databricks.com/company/careers/professional-services-operations/forward-deployed-engineer-8540455002)

- [Sr. Forward Deployed Engineer - National Security](https://databricks.com/company/careers/open-positions/job?gh_jid=8657468002)

- [Performance Tuning - Spark SQL documentation](https://spark.apache.org/docs/latest/sql-performance-tuning.html)

- [Web UI - Spark documentation](https://spark.apache.org/docs/latest/web-ui.html)

- [Row filters and column masks - Azure Databricks | Microsoft Learn](https://learn.microsoft.com/en-us/azure/databricks/data-governance/unity-catalog/filters-and-masks/)

- [Manually apply row filters and column masks - Azure Databricks | Microsoft Learn](https://learn.microsoft.com/en-us/azure/databricks/data-governance/unity-catalog/filters-and-masks/manually-apply)
