Snowflake World Tour 2026: AI agents are changing how we use data – but quality data still matters

Snowflake World Tour 2026 returned to Stockholm, bringing together more than 2,000 attendees to explore the latest developments in data and AI from Snowflake. The agenda featured more than 30 sessions and 67 speakers representing 11 different industries, covering everything from technical solutions and customer stories to Snowflake’s latest capabilities.

With such a broad agenda, it was easy to build the day around your own interests. My choices focused particularly on Snowflake’s AI capabilities, data governance and cost management, as well as how AI is changing data engineering.

Based on what we saw at the event, the direction is clear: we are moving beyond individual AI features towards agents and the concept of the Agentic Enterprise. AI is no longer limited to answering questions. It can use an organisation’s data, applications and business context to perform tasks.

At the same time, one message came through across almost every topic: the more responsibility we give to AI, the more important trusted data, governance and cost management become.

Snowflake World Tour 2026 Keynote – towards the Agentic Enterprise

The Agentic Enterprise was a central theme in Snowflake’s keynote. However, taking advantage of AI agents requires more than simply adopting a new AI service. Agents need trusted data, business context, governance and integrations with the organisation’s other systems.

Snowflake’s agentic environment is built around three key elements:

  • enterprise data and context,
  • AI models,
  • and software and applications.

These are brought together by an agent management layer designed to connect and manage the overall environment in a controlled way.

Examples highlighted in this context included Snowflake CoWork, which brings AI agents into employees’ daily workflows, and the more developer-focused Snowflake CoCo, which aims to accelerate the development of Snowflake-based software and data solutions with AI assistance.

At the same time, Snowflake emphasised the importance of an organisation’s own data, its business context and a governance model that enables agent activity to be controlled.

Another interesting announcement in the keynote was the collaboration between Snowflake and European cloud provider STACKIT. The collaboration aims to provide European organisations with a local cloud option. STACKIT operates data center in Germany and Austria.

Trusted data is still the foundation 

Despite the strong focus on AI, one familiar theme came up repeatedly throughout the day: without high-quality, well-governed data, there can be no trustworthy AI.

Swedbank’s presentation explored the bank’s Data & AI transformation and Snowflake’s role as its technical foundation.

The goal is to move from a situation where data is primarily seen as an IT responsibility towards a model where data is a shared resource across the entire business. Key objectives of the transformation include improving data quality and availability, reducing manual work, and bringing analytics and AI closer to everyday business decision-making.

Perhaps the most important message from the presentation was simple: trusted AI requires trusted data. Building AI solutions does not reduce the importance of a solid data platform, data quality or governance solutions. On the contrary, they become even more important.

More control over Snowflake governance and costs 

Snowflake has introduced several new capabilities for cost management and control. Costs can be monitored from the organisational level all the way down to individual users and teams. Budgets can also be defined for specific resources or based on tags.

AI-related costs can now be monitored more closely as well. This will become increasingly important as the use of AI services grows, as models and agents introduce a new cost dimension into the Snowflake environment. Per-user quotas and the ability to automate actions when budget thresholds are reached make these costs easier to manage.

The aim is to give users more freedom to build solutions while allowing the organisation to maintain control over costs.

Semantic Views create a shared meaning for data and AI 

As AI adoption grows, the importance of the semantic layer increases. If, for example, revenue is defined differently in Power BI, SQL queries and the context provided to an AI agent, the same question may produce several different answers.

Snowflaken Semantic View aims to address this by centrally defining business concepts, metrics, dimensions and the relationships between them in Snowflake. Semantic Studio makes semantic models easier to build. Previously, these models were largely created using YAML definitions, but they can increasingly be created and maintained visually and with the help of natural language.

Another interesting development is the integration of Snowflake’s semantic layer with other tools that use semantic models. For example, existing Power BI models can be used as a starting point for a Semantic View, while the same definitions can be utilised in Power BI and Excel.

Ultimately, the same business definitions can serve BI reporting, SQL queries, applications and AI agents.

Snowflake CoWork and Cortex Agents take AI from questions to action 

Snowflake CoWork was one of the most visible AI topics of the day. Its goal is to act as a personal work tool that understands an organisation’s data and context and can use different tools to perform tasks.

The key shift compared with earlier AI solutions is the move from a question-and-answer model towards action. Instead of simply analysing data and generating answers, an agent can use other agents, tools and enterprise systems to complete tasks.

This development is supported by capabilities such as CoWork automations, user-specific skills and integrations with other tools. In practice, the goal is that users do not need to know which agent or technical component is required for a particular task. Instead, CoWork can route the task to the appropriate agent.

AI is no longer simply a separate feature layered on top of the data platform. It is becoming an integral part of everyday work.

The role of the Data Engineer is changing with AI

Another topic that emerged throughout the day was the changing role of the Data Engineer. With AI, more tasks can be automated, including code generation, routine debugging, creating initial pipeline structures, and code migrations and refactoring.

This does not eliminate the need for Data Engineers. Instead, the focus is shifting towards understanding systems and solutions as a whole, designing data products, governance and understanding the business.

The value of a Data Engineer will therefore not be determined simply by how quickly they can make a pipeline or write SQL. The ability to understand what kind of solution should be built, how data should be governed and how the technical implementation serves the business will become increasingly important.

Snowflake World Tour 2026 – AI is moving from experimentation into the data platform

Snowflake World Tour was once again a valuable opportunity to see where Snowflake and the broader data industry are heading.

The future clearly looks increasingly agentic. At the same time, one of the strongest takeaways from the day was that AI does not reduce the importance of high-quality data, governance or cost management – quite the opposite.

The more opportunities we give AI to use enterprise data and perform tasks independently, the more important it becomes to understand what the data means, who is allowed to use it and at what cost.

-Asko Ovaska

Partner, Senior Data Engineer

Why you mustn’t settle on close enough numbers 

Sometimes I encounter attitudes such as “this KPI is in the ballpark” or “not matching exactly but in tolerance”. Would that suffice if the question was about the balance of your bank account, salary payment, or tax bill?

Oh but your inventory level or exact billing rate isn’t just that critical?

Well, if there’s a small recognizable error, how can you trust other numbers or maybe that same KPI but on different sub-selection?

Why “close enough” numbers are a problem

Remember that the law of big numbers conceals errors! Big aggregates converge to about right. That works in insurance pricing but not in discovering who’s working more productively than others or when evaluating quality of a product batch.

Perhaps the most important, jeopardized thing is trust in reports. As with trust in general, it’s easy to lose and next to impossible to regain. Do you really want to start your reporting renovation with just-almost-reliable numbers? Unfortunately, many do. Good luck leading change in your organization if the reports meant to support it have already given people reason not to trust them.

Way trickier situation is when wrong numbers are not recognized for being wrong. It’s not uncommon at all to discover later on that some calculation logic is not robust, and thus the business has been using wrong numbers for quite some time. It’s always a delicate issue to critically evaluate someone’s doings and faults, but it is something that has to be done, nevertheless. I bet many errors are ignored simply because of conformism and avoiding inconvenience.

This is not about predictions or otherwise inherently uncertain numbers but about numbers that can and should be logically and undisputably derivable. This is about erroneous logic, bad data, or both.

You can’t fix everything within a reasonable time frame and costs, but digging your head into the sand is a terrible choice.

Especially with ERP, CRM or otherwise human input data, there’s an endless stream of typing errors, missing data etc., but that is not actually bad data (even though I just used that phrase) but an important piece of reality.

Hand-in-hand with your KPI analytics endeavor, you should be getting tons of information about your data-generating processes. That is, about the ways Pekka and Tanja are filling values into those systems. Same but different goes for your IoT devices of course.

There are no shortcuts to data quality

I don’t have any magic bullets left in my clip, but I do have a few ordinary ones:

  • Verify results in reports vs. database vs. source systems 
  • You should practically always have row-level data as the basis in your Power BI data model so that you can verify numbers and utilize all the richness of data. If the volume is hundreds of millions of rows, there are ways for optimization, e.g. conditional aggregation. 
  • Don’t hide outliers, “duplicates” etc. lazily. Inspect their true origin and correct them at the source if possible. Document and set monitors for what you can’t fix completely. 
  • There’s no shortcut to quality and deep systemic understanding. I’m a humble student of Ferrari and Russo and their book “The Definitive Guide to DAX”. Ten years of solving problems with wits, YouTube and GPT doesn’t necessarily teach you about the whole system, or more importantly, what you don’t know. Think it this way: is it better to learn physics on your own or according to a high school curriculum? Very few of us are so passionately curious about this kind of subject that the first is a better option. The rest of us should learn the big picture from a curated comprehensive source. 

If you need analytics that are accurate, traceable and built to be trusted – rather than a collection of isolated solutions – let’s have a chat!

– Lauri Nuotio, Senior Analytics Engineer

Snowflake User Group Helsinki at Etlia – September 8, 2026

Etlia is proud to host and sponsor the September edition of Snowflake User Group Helsinki.

Join us for an evening dedicated to Snowflake, modern data platforms and analytics at Etlia’s office in Keilaniemi. The evening features a customer presentation from Fennia on identifying, tagging and utilizing PII data in Snowflake, followed by a recap of the most important announcements from Snowflake Summit 2026 presented by Snowflake. There will also be plenty of opportunities to network with members of the Finnish Snowflake community.

📅 Tuesday, September 8, 2026

🕔 5:00 PM–8:00 PM

📍 Etlia Oy
Finago Tower
Keilaniementie 1
02150 Espoo, Finland

🍕 Light bites and refreshments will be served.

🎟️ Participation is free, but registration is required.

👉 Register for the event

Agenda

5:00 PM
Arrivals, food & networking

5:30 PM
Snowflake User Group Helsinki welcome
Mika Heino & Asko Ovaska (Etlia)

5:40 PM
Jani Thind (Fennia)
Identifying, tagging and utilizing PII data in Snowflake
Presentation in Finnish + Q&A

6:30 PM
Break

6:40 PM
Elli Pennanen (Snowflake)
Snowflake Summit 2026 – Key announcements and their practical impact

7:30 PM
Networking

8:00 PM
Event ends

We look forward to welcoming the Finnish Snowflake community to Etlia!

SAP Sapphire Madrid 2026: From AI Assistance to the Autonomous Enterprise 

A Few Days in Madrid 

I had the opportunity to attend SAP Sapphire Madrid 2026 from 19 to 21 May, joining customers, partners, and SAP experts from across Europe to discuss the future of enterprise technology. While there were many interesting sessions throughout the event, the Global Keynote on 20 May, titled The Beginning of Better, stood out as the clearest expression of where SAP sees business technology heading next.

Presented by Christian Klein, Sebastian Steinhaeuser, Philipp Herzig, and Muhammad Alam, the keynote focused on a vision that goes beyond using AI as a productivity tool. Instead, SAP introduced its concept of the Autonomous Enterprise, where people define objectives, policies, and boundaries, while AI agents execute business processes within governed and trusted environments.

A Shift in How Enterprise AI Is Framed

One of my main observations from the keynote was that SAP is moving the conversation beyond AI assistants that simply help employees perform tasks faster. The focus is increasingly on AI systems that can carry out work autonomously while remaining under human oversight.

The centrepiece of this vision is the newly announced SAP Autonomous Suite. SAP described it as a framework that brings together AI agents, assistants, business applications, and governance capabilities into a single operating model. Rather than deploying isolated AI solutions, organizations would be able to orchestrate business processes that span multiple functions while maintaining visibility, compliance, and control.

This felt like an important distinction. Many organizations are currently experimenting with AI pilots and productivity tools, but the challenge often lies in connecting those capabilities to real business processes. SAP’s message was that enterprise AI needs to be embedded directly into operational systems rather than existing alongside them.

The Evolution of Joule

Another significant announcement was the evolution of Joule. Joule is now expanding into a broader AI workspace.

SAP presented Joule Work, Joule Assistants, Joule Agents, and Joule Studio 2.0 as components of a larger ecosystem designed to support both employees and autonomous business operations. The direction suggests a future where users interact with AI not only through conversational assistance, but also through specialized agents capable of executing defined business tasks.

What I found particularly interesting was how SAP positioned these capabilities as part of a coordinated environment rather than separate products. The emphasis was on enabling collaboration between people, AI assistants, and autonomous agents within existing business workflows.

Trust, Data, and Governance Take Center Stage

Perhaps the most important theme throughout Sapphire was trust.

SAP repeatedly emphasized that autonomous AI can only be effective when it is grounded in trusted business data and governed appropriately. This is where SAP Business Data Cloud, the new SAP Business AI Platform, and the SAP AI Agent Hub fit into the broader strategy.

The SAP Business AI Platform was presented as the foundation for building, governing, deploying, and operating enterprise AI agents. At the same time, SAP Business Data Cloud provides the business context that agents need to make informed decisions, while governance capabilities ensure transparency and accountability.

I was also interested to see SAP highlight its collaboration with Mistral AI and its commitment to trusted European AI and sovereign AI capabilities. For many organizations, particularly in Europe, questions around data sovereignty, regulatory compliance, and trusted AI infrastructure are becoming just as important as model performance.

Why This Matters

For organizations exploring AI adoption, the announcements at Sapphire were less about individual features and more about architecture and operating models.

Many companies have already discovered that deploying AI tools is relatively easy compared to scaling them across critical business processes. The challenge is creating an environment where AI can operate reliably, securely, and in alignment with business objectives.

The message from the keynote The Beginning of Better was clear: the next phase of enterprise AI is not simply AI assisted work. It is AI executed business processes operating under human governance. SAP’s vision of the Autonomous Enterprise brings together Joule, AI agents, trusted business data, and the SAP Autonomous Suite to make that transition possible.

Looking Ahead

Leaving Madrid, my impression was that enterprise AI discussions are becoming more practical and more operational. The focus has shifted away from experimentation toward questions of governance, execution, trust, and measurable business outcomes.

Whether organizations are ready for fully autonomous business processes today is another question. However, the direction presented at SAP Sapphire 2026 suggests that enterprise software is steadily evolving toward environments where humans set the goals and AI increasingly handles the execution. It will be fascinating to see how quickly that vision becomes reality over the coming years.

-Juuso Maijala, CEO

Leveraging AI in data stream loading: Matillion Data Productivity Cloud 

AI has been on everyone’s lips in the IT industry for the past few years, and its development has been rapid. Application providers have also developed and incorporated more and more AI-enabled capabilities into their software. Matillion is no exception. 

Matillion’s Data Productivity Cloud (DPC) includes several AI-based objects for use in developing loading pipelines. These include, among others, Copilot for natural language generation, AI Notes for automated documentation, and components for interacting with large language models (LLMs) such as OpenAI, Azure OpenAI, and Amazon Bedrock. In addition, Matillion DPC offers components for vector database operations (Pinecone, Snowflake) and for interacting with AI services such as Amazon Textract, Amazon Transcribe, and Azure Document Intelligence. 

In this blog, we will look at some of the opportunities offered by the Matillion Data Productivity Cloud (DPC) from three perspectives: 

  • Using AI in loading pipeline development with Matillion’s Maia AI assistant, 
  • Using AI in loading pipeline documentation, and 
  • Using Snowflake’s Cortex AI functions in a business use case. 

Using AI in Loading Pipeline Development 

In Matillion, loading pipelines have traditionally been built by dragging loading components onto a canvas, then connecting them and defining the desired processing rules. This has been a clear, intuitive, and easy way to build loading pipelines. Matillion’s AI assistant, Maia, takes this one step further. 

Maia is especially useful in cases where: 

  • You have only recently started using Matillion DPC or are otherwise inexperienced with ETL/ELT tools, 
  • The development environment is unfamiliar, or 
  • You want suggestions on how to solve a specific loading-related technical problem. 

In the example below, Maia was instructed to create a loading pipeline that: 

  • Loads customer and order data from the SNOWFLAKE_SAMPLE_DATA database, 
  • Joins the table data, and 
  • Creates a table and loads the data into the newly created table. 
Matillion can retrieve the correct source database and the desired tables from the metadata as separate table-read components, join the tables in a join component using keys, and use a rewrite component to create and load the table. 

Maia can be used for further development of the loading process or in individual loading components. In the example below, Maia is used to add metadata fields to the load and to perform calculations. 

Maia adds an Add Metadata calculation component to the loading pipeline, in which it includes the metadata fields and their functions. 
In the calculation component, the desired calculation is defined, and Maia creates the complete calculation formula. This is a useful feature, especially in cases where less commonly used functions are applied

Using AI in Loading Pipeline Documentation 

Documentation of loading pipelines often receives little attention or is not done at all. This frequently slows down further development of the loads, the fixing of errors, or, more generally, the understanding of the pipeline’s logic. Matillion DPC enables the automatic creation of descriptions for loading pipelines and individual components using AI, quickly and with just a few mouse clicks. 

In the example below, the Add Metadata calculation component contains added fields and calculations that we want to highlight. With Matillion, automatic description generation provides a clear picture of the operations performed within the object. 

AI-generated description of the Add Metadata component. 

By using Maia, it is possible to automatically create a complete description of the entire loading pipeline, including the loading steps and logic in detail. This feature enables effortless documentation of the entire pipeline and is highly beneficial when there is a need to understand the operation of a previously unfamiliar loading pipeline. 

Maia-generated summary of the loading pipeline. 

Snowflake Cortex AI Use Case 

Matillion DPC also enables the use of Snowflake’s AI capabilities with ready-made objects that can be added as part of regular loading pipelines. This low-code / no-code approach makes AI accessible to a wider group of users and lowers the threshold for adopting AI capabilities, as no special skills are required. 

Below is a simplified example where AI is used to create responses to customer feedback. The feedback has been received in three different languages (Finnish, Swedish, and English). The Matillion orchestration pipeline contains three different stages (transformation jobs), each using different Snowflake Cortex functions. The stages are: 

  • REVIEWS_1_TRANSLATE – Translates all feedback into English using the CORTEX.TRANSLATE function. 
  • REVIEWS_2_SENTIMENT – Determines the sentiment of the feedback (positive / negative) using the CORTEX.SENTIMENT function. 
  • REVIEWS_3_REPLY – Generates responses to the feedback using the CORTEX.COMPLETE function. 
Loading pipeline for automating customer feedback responses. 

In the first stage, the Cortex Translate object is used to translate customer feedback, that is loaded into Snowflake and written in multiple languages, into English. In the Cortex Translate object, the column to be translated is specified, along with the source language (in this case, automatic detection) and the target language, which is English. After the translation, the columns are renamed, and the data is loaded into the database for the next stage. 

Loading pipeline for translating customer feedback into English. 

In the second stage, the Cortex Sentiment object is used to identify the sentiment of the customer feedback. In the object’s settings, the column to which the sentiment analysis is applied is selected. Matillion creates a new column for the value, with a scale ranging from -1 (negative) to +1 (positive). Finally, in the loading pipeline, the necessary columns are renamed, and the data is loaded into the database for the final stage. 

Loading pipeline for determining the sentiment of customer feedback. 

In the final loading pipeline, the processing performed in the previous stages is used to generate a response to the customer feedback. At the start of the pipeline, a Filter object is used to split positive and negative feedback into separate data streams, based on the sentiment analysis carried out in the previous stage. Both data streams are then directed to their own Cortex Completions objects. The following are defined for these objects: 

  • The model to be used for generating the response. 
  • A system prompt that provides context for creating the response. 
  • A user prompt that specifies the actual response. 
  • The input to be used for generating the response: in this case, the feedback. 

After generating the response, the data streams are merged, transformed into a columnar format, the columns are renamed, and the data is loaded into the database, for example, to be used by customer service systems. 

Loading pipeline for generating customer feedback responses. 

Summary 

Matillion’s AI capabilities are not just technical accelerators, they are collaborative tools that bring data engineers, analysts, and business stakeholders onto the same page. They enable natural language interaction, automatic documentation generation, and the integration of diverse perspectives directly into data pipelines. Matillion bridges the gap between business and data systems, enabling understanding of actions without the need for technical skills such as code literacy. This opens the door to more agile development, faster iteration cycles, and, ultimately, more data products that are truly aligned with business priorities. 

The AI objects in Matillion DPC that can be used within loading flows speed up and simplify the adoption of these capabilities in organizations. The Snowflake Cortex objects are a good example of this. Matillion DPC is used to build and orchestrate the loading pipeline, while the actual data processing is performed in Snowflake. This eliminates data transfers, supports real-time execution, and ensures that security practices are followed consistently throughout. 

-Asko Ovaska

SAP and Databricks partner up—What’s in it for You?

Exploring the Future of Data Engineering with SAP Business Data Cloud

In today’s rapidly evolving digital landscape, businesses are constantly seeking innovative solutions to enhance their data management and analytics capabilities. On February 13, 2025, SAP launched the Business Data Cloud (BDC), a new Software-as-a-Service (SaaS) product designed to provide a unified platform for data and AI. According to SAP BDC is a comprehensive platform revolutionizing the way organizations handle data and artificial intelligence (AI) applications. In this blog post, I will delve into the key highlights of SAP Business Data Cloud and its collaboration with Databricks.

Introduction to SAP Business Data Cloud

SAP BDC combines several powerful components, including Datasphere, SAP Analytics Cloud (SAC), SAP BW, Databricks, and Joule (AI), to offer a comprehensive solution for data and AI needs. This integration offers AI capabilities, data management, and application support, making it ideal for businesses looking to fully utilize their data.

Key Features of SAP Business Data Cloud

The SAP Business Data Cloud is built to address a wide range of data and AI requirements. Some of its key features include:

  1. Comprehensive Data and AI Platform: BDC integrates various SAP and third-party data sources, providing a seamless flow from raw data to insightful analytics and AI applications. 
  1. Insight Apps: These ready-made SaaS products offer out-of-the-box solutions for data and AI needs, enabling businesses to quickly deploy and benefit from advanced analytics. 
  1. Custom Build Scenarios: BDC supports custom solutions, allowing organizations to combine SAP and third-party data to create tailored analytics and AI applications. It is also possible to copy Insight Apps components and to enhance copied functionalities with custom development.

The Role of SAP Databricks

A key feature of BDC is its integration with SAP Databricks. This collaboration brings Databricks’ powerful AI and machine learning (ML) functionalities to the SAP ecosystem, enabling businesses to leverage advanced analytics and AI capabilities within a single platform.

Benefits and Considerations for SAP BDC

The SAP Business Data Cloud offers several advantages that make it a compelling choice for businesses:

  1. Single SaaS Platform for Analytics, AI, and ML: BDC provides a unified platform that integrates various SAP and third-party data sources, enabling seamless data management and advanced analytics. 
  1. SAP Databricks AI/ML Functionalities: The integration with Databricks brings powerful AI and machine learning capabilities to the SAP ecosystem, enhancing the platform’s analytical capabilities. 
  1. Insights Apps: BDC includes ready-made SaaS products that offer out-of-the-box solutions for data and AI use cases, allowing businesses to quickly deploy and benefit from advanced analytics. 
  1. Tight Integration to Business Processes: BDC is designed to integrate seamlessly with existing business processes, ensuring that data analysis and AI applications are closely aligned with SAP business processes. 

While the SAP Business Data Cloud offers numerous benefits, there are a few considerations to keep in mind:

  1. A New Product – What is the maturity?: As a new product, businesses should evaluate the maturity of BDC for their respective use cases and consider any potential challenges during the implementation phase. 
  1. SAP Joule dependency – how to integrate into your overall architecture? Joule is yet another co-pilot AI interface to your stack. You have to make sure that for each use case there is well thought through user experience either through Joule or some other co-pilot integrating with Joule that is in line with your overall architecture e.g. MS co-pilot.  
  1. All in SAP or Use Both SAP BDC and Other Non-SAP Tools?: Organizations should still consider whether it makes sense to fully commit to the SAP ecosystem or to use a combination of SAP BDC and other non-SAP tools to meet their data and AI needs.

What to expect?

The SAP Business Data Cloud represents a significant leap forward in the realm of data engineering and AI. By combining the strengths of SAP’s data management tools with Databricks’ AI/ML capabilities, BDC offers a platform for businesses to enhance their data analytics and AI applications. As organizations continue to navigate the complexities of the digital age, solutions like BDC will play a crucial role in driving innovation and success. 

–Juuso Maijala, CEO & Founder

At Etlia Data Engineering we have a unique combination of expertise in both SAP and Databricks to support your business AI transformation. Want to know more? Book a meeting with us and let’s talk about how we can help your business to leverage SAP Business Data Cloud with Databricks!

Testing MS Fabric  Review on ”Auto-create report” -feature

One of our experts had previously produced a meaningful report of Finland’s Corona data. Now, after the launch of Microsoft’s new SaaS offering called Fabric, we will test its reporting feature which is supposed to ease the work of Data analysts and BI developers. With the “Auto-create report” -feature embedded to MS Fabric you can create insights from datasets with just one click. In the following text, we are going to compare the reports built by Fabric and our expert, and review if the quality of auto-created report matches the one produced by an expert. 

Fabric’s auto-created report of Finland’s Corona data 

How does it work? 

From Fabric UI you can conveniently access your data. You are able to create datasets from the files and tables you have uploaded to OneLake, which is this new unified data source we discussed in the previous blog concerning MS Fabric. By selecting the dataset you would like to create the report of, you can decide whether you like to build the report from scratch or if you want the Fabric to automatically build the report. 

When you decide to automatically create the report, Fabric picks up the columns from the tables it thinks are the most meaningful and creates the visuals to reflect the insights of that data. It creates a quick summary page to show the most important highlights on its opinion. It also writes a short text to summarize the insights of the visuals. You can then change the data you want to be projected and it automatically builds new visuals of the selected data. 

Comparison 

Using the “Auto-create report” -feature you can easily build a sufficient report which tells effectively the key insights of the data. You probably still need to do some work selecting the right data to be projected, because it doesn’t necessarily pick up the right columns right away. The report it creates, may be good enough, if you just need to quickly check what is happening. However, the report it creates, isn’t visually as exquisite or informative as one created by an expert. Also, it only offers a quick summary of the data, whereas a human can create multi page report offering deep understanding of the matter. You can also change the type of the visualizations in the automatically created report, but it is as simple to build the report from scratch. If you want to build a presentable report with the help of “auto-create report” -feature, you have to put as much thought and effort on it as if you were to build the whole thing from scratch. 

In conclusion, we think that this feature is nice add to Power BI, because anyone can easily check the insights that the data has to offer and make decisions based on that information. Anyway, if you want to create a report that offers powerful support for your presentation, you still need to use some time on building the report and empathizing the major data points. 

Future of AI Analytics 

Even though the quality of the automatically created report isn’t yet quite as insightful as the report created by human expert, it still is impressive, how well it can connect different types of data and produce meaningful visuals all by itself. AI and machine learning technologies have been rapidly evolving in recent years and Data Analytics offers great usage for those. They are already great at identyfying patterns and analyzing the relationships and dependencies between variables. Therefore, we believe that there is still room for this “auto-create report” -feature to improve. In the future, it might be able to interpret and communicate the information hidden in the data even better than the brightest expert. 

At the moment, the trend seems to be that we are trying to advantage AI by using generative AI language models as a trusted helper that will do the hand-on work for us. Microsoft has informed us about the copilot feature which will be included in the Fabric offering but isn’t available yet in the public preview version. They have showed us how you’ll be able chat and tell what information you want to know from the data. It can create measures and SQL views. Of course, it can create visuals, but it can answer more sophisticated questions too. For example, it can show you with visuals the reasons why something has happened or give suggestions via chat on how you could improve certain values. With the copilot, the only thing left for humans to do, is to know what to ask. Often those questions repeat themselves so maybe we might be able to automate also that task someday. 

Etlia Data Engineering announces completion of personnel share offering

Etlia Ltd

News release

28 April 2023 – 09:00 EET

Etlia Data Engineering has today closed it’s first personnel share offering. All Etlia’s employees participated in the offering with full subscription rights making all the employees also shareholders of the Company.

“Our personnel offering was 100% success! I am thrilled to see such engagement and interest into our share offering. I am proud that now all our employees are also Etlia’s shareholders. Our intention is to continue personnel offerings also in the coming years alongside our partner program which was launched this year.” says Juuso Maijala, CEO.

“It is fantastic to see the huge enthusiasm of Etlians and their commitment into company’s growth journey. Using ITA66a§ (Finnish: TVL66a§) framework provides an excellent way to engage personnel and I can recommend it to any company seeking to boost it’s growth through a share based incentive program.“ says Mikko Koljonen, Board Member.  

Additional information:

Juuso Maijala, CEO

juuso.maijala@etlia.fi

+358 50 532 0157

Mikko Koljonen, Board Member

mikko.koljonen@etlia.fi

+358 50 36 28 218

Etlia Ltd shortly:

Etlia is a data engineering company.

We help our customers create business value from data by leveraging major business process platforms and external sources. We offer top experts the best platform and community to grow professionally. Our company was founded in 2013. We are based in Espoo, Finland.

Synapse vs Databricks: A Comparison 

From Databricks to Synapse: A Data Architect’s Journey 

As a Data Platform Architect/ Engineer working with several clients in Finland, I have extensive experience using Azure Databricks and Azure Data Factory (for notebook orchestration). Recently, however, one of my clients made the decision to switch to Azure Synapse Analytics. In this post, I will share my journey of transitioning from Databricks to Synapse and provide insights that may help you make a more informed decision if you are considering either of these platforms. 

When it comes to choosing between Synapse and Databricks for your data processing needs, there are several factors to consider. Firstly, we will take a closer look at some of the key features of each platform and then finally my opinion on the matter.

Data Storage, Resource Access, and DevOps Integration 

When comparing Databricks and Synapse, it is important to consider the availability of certain features. For example, Databricks allows you to use multiple notebooks within the same session – a feature that is not currently available in Synapse. Another key difference between the two platforms is the way they handle data storage. Databricks provides a static mount path for your storage accounts, making it easy to navigate through your data like a traditional filesystem. In contrast, Synapse requires you to provide a ‘job id’ when reading data from a mount – an id that changes every time a new job is run. 

When it comes to accessing resources, Synapse offers linked service access management – a feature that allows for cleaner and more manageable connections between different services via Azure. In contrast, Databricks relies on tokens generated by service principals for resource access. However, Databricks does have an advantage when it comes to bootup time – boasting faster speeds than Synapse. On the other hand, Synapse has better DevOps integration compared to Databricks. 

Features, Performance and Use Cases 

There are several other key differences between Databricks and Synapse that are worth considering. For example, Databricks currently offers more features and better performance optimizations than Synapse. However, for data platforms that primarily use SQL and have few Spark use cases, Synapse Analytics may be the better choice. Synapse has an open-source version of Spark with built-in support for .NET applications, while Databricks has an optimized version of Spark that offers increased performance. Additionally, Databricks allows users to select GPU-enabled clusters for faster data processing and higher concurrency. 

User Experience 

In terms of user experience, Synapse has a traditional SQL engine that may feel more familiar to BI developers. It also has a Spark engine for use by data scientists and analysts. In contrast, Databricks is a Spark-based notebook tool with a focus on Spark functionality. Synapse currently only offers hive metadata GUI but with Unity Catalog, Databricks takes it to another level of creating the metadata hierarchy. 

Managing Workflows with External Orchestration Tools 

One important aspect to understand when using notebooks in Databricks is the lack of an in-built orchestration tool or service. While it is possible to schedule jobs in Databricks, the functionality is quite basic. For this reason, in many projects we used Azure Data Factory to orchestrate Databricks notebooks. In a recent Databricks meetup, one participant mentioned using Apache Airflow for orchestration on AWS – though I am not sure about GCP. This is a crucial point to consider because Synapse bundles everything under one umbrella for seamless integration. Until Databricks produces an alternative solution, you will need to use it alongside ADF (Azure Data Factory) or Synapse for orchestration.  

Feature Databricks Azure Synapse Analytics 
Multiple notebooks within same session Yes No 
Data storage handling Static mount path for storage accounts Requires ‘job id’ when reading data from a mount 
Resource access management Tokens generated by service principals Linked Service access management 
Bootup time Faster speeds than Synapse Slower speeds than Databricks 
DevOps integration Less integration compared to Synapse Better integration compared to Databricks 
Features and performance optimizations More features and better performance optimizations than Synapse Fewer features and less performance optimizations than Databricks 
SQL support Less support for SQL use cases Better support for SQL use cases 
Spark engine Optimized version of Spark that offers increased performance Open-source version of Spark with built-in support for .NET applications 
GPU-enabled clusters Allows users to select GPU-enabled clusters for faster data processing and higher concurrency Not available in Synapse now. 
User experience Spark-based notebook tool with a focus on Spark functionality Traditional SQL engine that may feel more familiar to BI developers. Also has a Spark engine for use by data scientists and analysts.  
Real-time Co-Authoring Databricks Notebooks has as real-time co-authoring (both authors see the changes in real-time) Synapse Notebooks has co-authoring of Notebooks, but one person needs to save the Notebook before another person sees the change 
Orchestration tool or service Lacks an in-built orchestration tool or service. Needs to be used alongside ADF or Synapse for orchestration. Bundles everything under one umbrella for seamless integration. 
Synapse vs Databricks feature comparison summary table. 

Choosing Between Databricks and Synapse: Which One Is Right for You? 

Ultimately, the choice between these two platforms will depend on your specific needs and priorities. Nah! I will not leave you with a diplomatic answer. In my opinion (could be controversial based on your cloud bias and when are you reading this) if your infra is on AWS/GCP, your priority is data processing efficiency and access to latest spark and delta features go for Databricks. 

On the other hand, if your infrastructure is primarily based on Azure and your use case involves data preparation for a data platform with data modeling on a Datalake (reach out if you are interested to know how), then Azure Synapse may be the better choice. Synapse has more features in development for future releases – something that has not been announced by Databricks yet. Good luck! And stay tuned for upcoming series focusing on ML, streaming, delta and partitioning. 

.