What is our primary use case?
Azure Data Factory's main use case is to take raw data from the source and move it to the destination like ADLS Gen2. I build end-to-end pipelines with backfilling features in Azure Data Factory. Along with that, I also schedule the pipeline.
I have used Azure Data Factory for copying raw data from a source to a final destination by using drag-and-drop features in a pipeline, as well as for scheduling and orchestration of the pipeline. These are the main features that I have used in Azure Data Factory.
What is most valuable?
The best features about Azure Data Factory are that it gives the features of drag-and-drop activities. I don't have to write each line for different activities such as copy activity or web activities. I just have to drag and drop those activities and make a pipeline. This is one of the main features. The second feature is that I can schedule my pipeline. Additionally, I can orchestrate my pipelines. If I have a lot of pipelines for a particular job, I need to do the orchestration. Another feature is that it is used to copy the raw data from a source to a destination. These are the main features that Azure Data Factory provides.
Apart from that, as I mentioned, Azure Data Factory is also used for backfilling features. If in my pipeline I am running my pipeline and I come to know that for a particular date my data has not come, I can add the backfilling features in my pipeline. Also, if I have already copied my raw data to the destination and I want to run that pipeline again, I should not have the data that I have already copied come to my destination again. For that, I can use the backfilling features. These are the features I would like to highlight.
Since I have mentioned that I can use the drag-and-drop activities, by using that, Azure Data Factory is saving a lot of time. My organization has already been saving the time and generating more income with that. Second is that I can use the AI features as well in Azure Data Factory because of its built-in AI. These are some of the features in which my organization is, of course, making a profit.
If I were to create a pipeline by writing everything line by line, it would take almost seven to eight days to complete a pipeline. But using the drag-and-drop activities, I can do it in minutes, around 15 to 20 minutes. It is saving almost five to six days of time. Taking a rough example of 6 into 5, it is saving 30 hours of time.
What needs improvement?
If the AI features were more improved so that I don't have to provide each and every detail, Azure Data Factory could be improved in a much better way by improving the AI features.
For example, if I want to fetch any data from a raw source, I need to provide each and every detail. But if I am just uploading my raw data and if AI will sync with that data, it can analyze that data and give me proper suggestions on how that should be done in a proper way. Automatic suggestions could improve in a much better way.
As I have mentioned, the AI features as well as more drag-and-drop activities could be improved. If I am making a pipeline, it should give me suggestions, such as which activity should be used, so that I don't have to remember each activity. If I have used one activity, I shouldn't have to remember what activity should I use next. It should give auto-suggestions. That is why I have given a nine out of 10.
Currently, I don't know about its governance and security, but in view of its improvement, I think Azure Data Factory should improve in these areas.
As I already mentioned, the AI features should be improved. Also, the auto-suggestion features should also improve.
What do I think about the stability of the solution?
In terms of accuracy and reliability of its output, it is good. It gives me the accurate result. Reliability is also good. I can rely on it. It is good overall.
What other advice do I have?
Azure Data Factory is good, especially for a beginner because when I was a beginner, I started my journey of data engineering with Azure Data Factory only. When I was learning it, using the drag-and-drop activity made it very simple to learn for a beginner. Additionally, with the help of the AI features, I was testing a lot of things. Also, rescheduling the pipeline and orchestrating the pipeline make it good for a data engineering professional. I gave this product a rating of 9 out of 10.