- Records: parses CSV and Excel (
.xlsx,.xls) files into tables, with automatic schema detection and type inference. - Unstructured: uploads files of any type (PDF, Word, images, and so on) to a Nekt volume and creates a metadata table describing them.
Configuring SFTP as a Source
In the Sources tab, click on the “Add source” button located on the top right of your screen. Then, select the SFTP option from the list of connectors. Click Next and you’ll be prompted to add your access.1. Add account access
Configure your server connection details:- Protocol: How Nekt reaches your server. Choose SFTP for file transfer over SSH, FTPS for FTP encrypted with TLS, or FTP for a plain, unencrypted connection. Defaults to SFTP.
- Hostname: The host for accessing your server.
- Port: The port for accessing your server. Usually
22for SFTP and21for FTP and FTPS. - Username: The username to access your server.
- Password: The password to access your server. For SFTP, use either password or private key.
- Private key: The private key for accessing your SFTP server. Use either password or private key. Not used by FTP and FTPS.
- Private key passphrase: The passphrase that decrypts your private key. Leave it empty if the key was generated without a passphrase. Not used by FTP and FTPS.
- Folder path: Path of the folder on the server where the files are located.
- Mode: Choose Records to parse CSV and Excel files into tables, or Unstructured to upload raw files to a Nekt volume.
- Files volume: Select the Nekt volume where the fetched files will be uploaded. This is required when the mode is set to Unstructured.
- Consider single stream: (Records mode) Consider all files in the folder as belonging to the same stream. When enabled, the matched CSV files are merged into a single
datastream and the matched Excel files into a singledata_excelstream, instead of one stream per file. - Delimiter: (Records mode, CSV files) Character that separates columns in the CSV file.
- Sheet Name: (Records mode, Excel files) Name of the sheet to read from Excel files. Defaults to the first (active) sheet if not provided.
- Pattern to match files: A regular expression matched against file names. In Records mode it applies when all files are considered a single stream; in Unstructured mode it filters which files are uploaded to the volume. Leave empty to include every file.
Some servers only accept encrypted connections. If a connection fails on FTP, try FTPS with the same host, port and credentials.
2. Select streams
In the discovery process, the connector lists what it found on the server.- In Records mode: The connector automatically detects CSV (
.csv) and Excel (.xls,.xlsx) files sitting directly in your specified folder path and maps each one to a stream. It’s important to ensure these files are properly formatted in a tabular way, so the data can be mapped correctly to a table in your Lakehouse. - In Unstructured mode: The connector lists a single stream called
sftp_filesthat replicates the raw files to your configured Nekt volume and creates a metadata table. Unlike Records mode, this walks the folder and all of its subfolders.
Tip: The stream can be found more easily by typing its name.Select the streams and click Next.
3. Configure data streams
Customize how you want your data to appear in your catalog. Select the desired layer where the data will be placed, a folder to organize it inside the layer, a name for each table (which will effectively contain the fetched data) and the type of sync.- Layer: choose between the existing layers on your catalog. This is where you will find your new extracted tables as the extraction runs successfully.
- Folder: a folder can be created inside the selected layer to group all tables being created from this new data source.
- Table name: we suggest a name, but feel free to customize it. You have the option to add a prefix to all tables at once and make this process faster!
- Sync Type: in Records mode the syncs will always be Full Sync. In Unstructured mode the
sftp_filesstream is Incremental, using each file’s modification date on the server, so every run only transfers files that were added or changed since the last one. Read more about sync types here.
4. Configure data source
Describe your data source for easy identification within your organization, not exceeding 140 characters. To define your Trigger, consider how often you want data to be extracted from this source. This decision usually depends on how frequently you need the new table data updated (every day, once a week, or only at specific times). Once you are ready, click Next to finalize the setup.5. Check your new source
You can view your new source on the Sources page. If needed, manually trigger the source extraction by clicking on the arrow button. Once executed, your data will appear in your Catalog.Streams and Fields
Depending on the selected mode, the connector exposes different streams:- Records Mode: The connector dynamically generates streams based on the CSV and Excel files found in the specified folder path. The schema of each stream reflects the columns inside that specific file. If Consider single stream is enabled, all matched files are merged into one stream per file type (
datafor CSV,data_excelfor Excel). - Unstructured Mode: The connector provides a single stream containing metadata for the files it uploads to your Nekt volume.
Files (sftp_files)
Files (sftp_files)
Stream emitted in unstructured mode. Walks the configured folder and all of its subfolders, uploads every file to a Nekt volume, and emits one metadata record per file. Downstream targets will materialize a single
sftp_files table.Key fields:Notes:
- Files of any type are uploaded, not just CSV and Excel. Use the Pattern to match files setting to restrict which ones are picked up.
- Because the volume identifies files by name, a file’s folder path is folded into the name it is stored under, so two files with the same name in different folders never overwrite each other.
CSV Files
CSV Files
Stream generated in records mode from
.csv files found on the server.Notes:- The schema and fields are dynamically discovered based on the header row of the CSV files.
- Fields are parsed according to their inferred types (string, integer, float, boolean, etc.).
- If a specific delimiter is used, it should be configured in the source settings.
Excel Files
Excel Files
Stream generated in records mode from
.xls and .xlsx files found on the server.Notes:- The schema and fields are dynamically discovered based on the header row of the specified sheet.
- If no sheet name is provided in the configuration, the connector defaults to reading the first available sheet in the workbook.
- Fields are parsed according to their inferred types (string, number, boolean) based on sample rows.
- Column headers are sanitized automatically to ensure they are safe for downstream tables.