Posts

Showing posts with the label Data warehouse

Create Bonze-Silver-Gold-table in minutes in Databricks following mediation architecture

Image
 AI is doing a good job in Databricks platform.  Earlier, following typical mediation architecture, we used to write bunch of code to read the raw data, create a bronze table and then silver table, and finally the gold table.  Now, I open databricks and upload the raw CSV data file. Next, I prompt properly and ask AI to create the code for me.   Code generate within seconds. Ready to execute. However, AI is not always right as I'm. This is acceptable. Code unable to understad the final level aggregration columns there. Good news is AI is there to diagnose the error and recomend the right code. AI, understand what code change is required while doing analysis generated SCHEMA in the dataframe. AI recomending you the right python code to fix this. Let's divide the phase in two sefment. Creating Bronze and Silver table in one segment and Gold one is another code segment. Once the Silver DataFrame generates, print the dataframe schema and review the column before wri...

How to bring on-premise SQL Server data into Azure Fabric Lakehouse

Image
On-premise data gateway - one of the important feature/service Azure Fabric platform is using to bring your on-premise data into Azure cloud. The data movement is secure and encrypted. You can review my my earlier post  https://splaha.blogspot.com/2025/03/use-on-premise-gateway-in-fabric-and.html   where I explain how to configure On-premise data gateway using Step 1 and Step 2. In this exaple, I'm going to use the same data conenction to pull my on-premise SQL tables using custom join SQL into Lakehouse. The entire post is self explanatory with step-by-step snapshot. I hope this will help you to understand each and individual steps as well as to execute the same at your end. Let's start with my on-premise SQL server details. As you can see the below picture, it talks about the database Employee where there are two tables EMP and DEPT. I'm going to use the same SQL into Fabric to pull the data. I assume you already created the on-premise data gateway using my earlier post...

Use On-Premise-Gateway In Fabric and copy data from On-prem to Azure Fabric

Image
On-Premise data gateway is an application which is required to be installed in your laptop/computer/server which connect to cloud and uses cloud services to send data back-and-forth from on-prem to cloud. This is secure data connection and data encryption is in place. Here is more about on-premise data gateway - https://learn.microsoft.com/en-us/data-integration/gateway/service-gateway-onprem?toc=%2Ffabric%2Fdata-factory%2Ftoc.json https://learn.microsoft.com/en-us/data-integration/gateway/service-gateway-onprem?toc=%2Ffabric%2Fdata-factory%2Ftoc.json Architecture of it - https://learn.microsoft.com/en-us/data-integration/gateway/service-gateway-onprem-indepth?toc=%2Ffabric%2Fdata-factory%2Ftoc.json https://learn.microsoft.com/en-us/data-integration/gateway/service-gateway-onprem-indepth?toc=%2Ffabric%2Fdata-factory%2Ftoc.json Step 1 First thing first - You need to install on-premise data gateway in your lcoal. Second thing is to open the on-premise data gateway and ensure th...

How to fix ModuleNotFoundError - No module named pymongo in Notebook

Image
One common issues while working with python or spark is getting error message which says - module not found. Module not found here points that proper configuration or installation is required in respect to libray level. Once this is done, application will able to find my required module. For example, while using below code in Fabric Notebook and trying to connect with MongoDB database and display records, getting no module found error message. Running the code, it is throwing error - ModuleNotFoundError - No module named pymongo To fix the abobe issue, it is required to install the required libries. #install the required packages ! pip install pymongo ! pip install certifi Once done, Azure notebook now able to connect with Mongo database, and getting the confirmation log from Azure. Collecting pymongo Downloading pymongo-4.6.3-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (676 kB) ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 676.9/676.9 kB 17.3 MB/s eta 0:00:00 a 0:00...

Microsoft Fabric - Use Lakehouse to upload source

Image
Microsoft Fabric comes up with multiple capabilities/wings and one of it is Data Enginnering where you brings your data to next generation AI. In Data Enginnering platform, you are going to load your data, perform operation on your data to process it, and finally display your finetune data in nice way using Power BI capabilities. Tables are Files hold your data in Data Enginnering landscape.  Table will allow to hold data in table structure format while you can upload your data file using csv/json/parquet. Upload option is there to upload your file(s) into Microsoft Fabric.              Files uploading in Lakehouse Files uploaded in Lakehouse Simple way to display record is to create one Notebook and drag the file there, Fabric will create the code for you :) Click on the table data, and options are there to display the data in different format. For example, you can view your data in chart format (like bar/chart/pie).   Bar Format Pie Format...