1
00:00:00,005 --> 00:00:02,008
- [Instructor] Technology validation requires

2
00:00:02,008 --> 00:00:05,005
developing some functionality that cuts through

3
00:00:05,005 --> 00:00:07,009
all layers of technology stack.

4
00:00:07,009 --> 00:00:10,000
User interface at front end,

5
00:00:10,000 --> 00:00:12,004
business logic in middle layer,

6
00:00:12,004 --> 00:00:15,000
and data at the back end.

7
00:00:15,000 --> 00:00:19,000
So we need to build a proof of concept.

8
00:00:19,000 --> 00:00:22,002
And to do that, let us set up the stack.

9
00:00:22,002 --> 00:00:24,000
This is a deployment diagram

10
00:00:24,000 --> 00:00:26,007
that shows various layers in the stack.

11
00:00:26,007 --> 00:00:29,000
It helps to have some visual layout

12
00:00:29,000 --> 00:00:32,002
before we go ahead and set up the environment.

13
00:00:32,002 --> 00:00:34,001
Starting from the user's end,

14
00:00:34,001 --> 00:00:37,002
the diagram shows browser on the user device

15
00:00:37,002 --> 00:00:41,006
with HTML, CSS, and JavaScript at the front end.

16
00:00:41,006 --> 00:00:45,005
Then I plan to use Apache Tomcat as my web server

17
00:00:45,005 --> 00:00:47,007
that will host my application.

18
00:00:47,007 --> 00:00:50,008
The application will be in the form of a web archive

19
00:00:50,008 --> 00:00:53,004
that is a war file.

20
00:00:53,004 --> 00:00:58,006
The database called a Red30DB will be on MySQL

21
00:00:58,006 --> 00:01:00,003
which will have some connection

22
00:01:00,003 --> 00:01:03,000
with the external USDA database.

23
00:01:03,000 --> 00:01:06,000
This is the architecture of my production environment.

24
00:01:06,000 --> 00:01:10,005
Now we need to think about the dev environment as well.

25
00:01:10,005 --> 00:01:12,003
In agile methodology,

26
00:01:12,003 --> 00:01:15,005
the entire cycle from dev to test,

27
00:01:15,005 --> 00:01:18,005
to stage, to production needs to flow

28
00:01:18,005 --> 00:01:21,003
in one seamless pipeline.

29
00:01:21,003 --> 00:01:23,008
DevOps, which is the agile approach

30
00:01:23,008 --> 00:01:28,004
for software development life cycle makes that happen.

31
00:01:28,004 --> 00:01:32,006
When you integrate dev and integration testing environment,

32
00:01:32,006 --> 00:01:35,003
you achieve continuous integration.

33
00:01:35,003 --> 00:01:37,006
When you are able to promote your code

34
00:01:37,006 --> 00:01:40,003
to staging environment incrementally

35
00:01:40,003 --> 00:01:44,000
where your end users can do some acceptance testing,

36
00:01:44,000 --> 00:01:46,001
you'll get continuous delivery.

37
00:01:46,001 --> 00:01:48,005
And when you are able to move your code

38
00:01:48,005 --> 00:01:50,002
incrementally to production,

39
00:01:50,002 --> 00:01:52,004
then you have continuous deployment.

40
00:01:52,004 --> 00:01:55,004
Now to get this level of sophistication

41
00:01:55,004 --> 00:02:00,009
requires building a fairly sophisticated DevOps tool chain.

42
00:02:00,009 --> 00:02:04,001
In this case study, we will limit ourselves

43
00:02:04,001 --> 00:02:09,006
to using JUnit for unit testing, Maven for building,

44
00:02:09,006 --> 00:02:12,002
GitHub for source code management,

45
00:02:12,002 --> 00:02:15,009
and Jenkins for promoting our code to production.

46
00:02:15,009 --> 00:02:19,009
Let us take a look at how it is all set up.

47
00:02:19,009 --> 00:02:25,001
The developer works on dev machine using an IDE,

48
00:02:25,001 --> 00:02:27,006
which is in this case is eclipse.

49
00:02:27,006 --> 00:02:30,003
It has JUnit for unit testing,

50
00:02:30,003 --> 00:02:32,009
Maven for build automation,

51
00:02:32,009 --> 00:02:36,009
and Git for staging configured within it.

52
00:02:36,009 --> 00:02:40,002
I have the Tomcat for dev environment

53
00:02:40,002 --> 00:02:42,003
also configured on this machine

54
00:02:42,003 --> 00:02:46,008
named as TomcatDev on port 8082.

55
00:02:46,008 --> 00:02:51,006
The code deployed on TomcatDev connects to MySQL server

56
00:02:51,006 --> 00:02:54,002
also running on the dev machine.

57
00:02:54,002 --> 00:02:56,006
As code is developed and tested,

58
00:02:56,006 --> 00:02:58,009
it is staged using git staging

59
00:02:58,009 --> 00:03:02,004
and then pushed onto the git repo.

60
00:03:02,004 --> 00:03:07,001
I have Jenkins hosted on my localhost at port 8080.

61
00:03:07,001 --> 00:03:10,004
That will pull the code from git repo,

62
00:03:10,004 --> 00:03:14,000
use Maven plugin to build a war file,

63
00:03:14,000 --> 00:03:16,002
and use the deployment plugin

64
00:03:16,002 --> 00:03:18,007
to deploy onto the production server.

65
00:03:18,007 --> 00:03:22,002
The production server is on AWS Cloud

66
00:03:22,002 --> 00:03:24,005
hosted on an EC2 instance

67
00:03:24,005 --> 00:03:28,001
that has another Tomcat server running on it.

68
00:03:28,001 --> 00:03:31,008
When Jenkins deploys a war file on production server,

69
00:03:31,008 --> 00:03:35,005
it connections with production database, red30db,

70
00:03:35,005 --> 00:03:39,000
which is running on Amazon Cloud Relation Database Service

71
00:03:39,000 --> 00:03:41,004
or RDS.

72
00:03:41,004 --> 00:03:45,008
As you can see, I have not shown the USDA database here,

73
00:03:45,008 --> 00:03:49,003
and that is because the data coming from USDA

74
00:03:49,003 --> 00:03:51,003
is static and read-only.

75
00:03:51,003 --> 00:03:54,009
So instead of configuring hard connection to it,

76
00:03:54,009 --> 00:03:57,007
I'll simply put the data provided by USDA

77
00:03:57,007 --> 00:04:00,006
into my database, red30db.

78
00:04:00,006 --> 00:04:02,005
Now there is just one thing missing,

79
00:04:02,005 --> 00:04:04,006
and that is the data model.

80
00:04:04,006 --> 00:04:07,000
At this point, I will encourage you

81
00:04:07,000 --> 00:04:12,001
to visit USDA FoodData Central at the URL shown here.

82
00:04:12,001 --> 00:04:14,007
The data we want to use in this project

83
00:04:14,007 --> 00:04:19,002
is available as a CSV or Access database,

84
00:04:19,002 --> 00:04:21,001
along with some documentation.

85
00:04:21,001 --> 00:04:23,004
We do not need all that data.

86
00:04:23,004 --> 00:04:25,007
For the use cases we have identified,

87
00:04:25,007 --> 00:04:28,005
we need just four tables to start with,

88
00:04:28,005 --> 00:04:32,005
and these tables are food, branded food,

89
00:04:32,005 --> 00:04:35,000
nutrient, and food nutrient.

90
00:04:35,000 --> 00:04:37,001
I have drawn the entity relationship

91
00:04:37,001 --> 00:04:40,006
that has ER model here, showing key entities,

92
00:04:40,006 --> 00:04:42,007
their primary and foreign keys,

93
00:04:42,007 --> 00:04:44,007
data type for each field,

94
00:04:44,007 --> 00:04:48,003
relationships, and their cardinalities.

95
00:04:48,003 --> 00:04:50,005
These tables have many more columns,

96
00:04:50,005 --> 00:04:52,002
but I have not modeled them here

97
00:04:52,002 --> 00:04:54,006
as they're not relevant for us.

98
00:04:54,006 --> 00:04:57,000
Next step, let us take a look at the data

99
00:04:57,000 --> 00:05:02,000
on SQL server to understand its basic structure and values.

100
00:05:02,000 --> 00:05:04,000
I have already migrated the data

101
00:05:04,000 --> 00:05:08,009
for these four tables from Access to red30db in MySQL.

102
00:05:08,009 --> 00:05:12,007
I'm using MySQL Workbench to view the data.

103
00:05:12,007 --> 00:05:15,008
And so let us start the default table.

104
00:05:15,008 --> 00:05:19,009
I have already made red30db as my default schema.

105
00:05:19,009 --> 00:05:25,006
So now I can say select start from food,

106
00:05:25,006 --> 00:05:28,001
and then run the query.

107
00:05:28,001 --> 00:05:32,000
And this shows FDC ID, which is the primary key,

108
00:05:32,000 --> 00:05:34,006
and description, which has the product menu.

109
00:05:34,006 --> 00:05:36,003
It has some other columns,

110
00:05:36,003 --> 00:05:39,004
which are not relevant for us at this stage.

111
00:05:39,004 --> 00:05:43,004
The next table is branded food.

112
00:05:43,004 --> 00:05:45,009
This table will show some additional information

113
00:05:45,009 --> 00:05:47,008
about the food products,

114
00:05:47,008 --> 00:05:49,007
starting with the primary key once again,

115
00:05:49,007 --> 00:05:54,001
FDC ID, the brand owner, ingredients,

116
00:05:54,001 --> 00:05:57,000
serving size, serving size unit,

117
00:05:57,000 --> 00:06:00,003
household survey, and some other columns

118
00:06:00,003 --> 00:06:04,000
that are not relevant for us at this stage.

119
00:06:04,000 --> 00:06:08,004
Third is nutrient table.

120
00:06:08,004 --> 00:06:10,000
This is a small table.

121
00:06:10,000 --> 00:06:15,007
It has nutrient ID, name, and measuring unit

122
00:06:15,007 --> 00:06:17,008
with some other data.

123
00:06:17,008 --> 00:06:21,005
And finally, food nutrient,

124
00:06:21,005 --> 00:06:26,007
which has all products and the nutrients in each product.

125
00:06:26,007 --> 00:06:28,000
This is the largest table,

126
00:06:28,000 --> 00:06:30,003
and it has maximum number of rows.

127
00:06:30,003 --> 00:06:34,005
So as you can see here, food has close to 300,000,

128
00:06:34,005 --> 00:06:38,002
branded food has slightly lesser number of rows,

129
00:06:38,002 --> 00:06:41,000
nutrient had just about 220,

130
00:06:41,000 --> 00:06:43,002
and now we can see food nutrient,

131
00:06:43,002 --> 00:06:45,009
it has close to half a million rows.

132
00:06:45,009 --> 00:06:48,002
Here you can see the FDC ID,

133
00:06:48,002 --> 00:06:50,006
which is the product identifier,

134
00:06:50,006 --> 00:06:53,006
nutrient ID, which is the nutrient identifier,

135
00:06:53,006 --> 00:06:56,004
and the amount which tells us how much nutrient

136
00:06:56,004 --> 00:06:59,000
is there in that product.

137
00:06:59,000 --> 00:07:03,000
Now to see all these four tables coming together,

138
00:07:03,000 --> 00:07:05,008
I have written one simple query,

139
00:07:05,008 --> 00:07:09,005
which shows information for a particular product

140
00:07:09,005 --> 00:07:13,006
with FDC ID S344604.

141
00:07:13,006 --> 00:07:15,004
And if I run this query,

142
00:07:15,004 --> 00:07:21,006
I can see the product FDC ID, description, brand owner,

143
00:07:21,006 --> 00:07:24,001
nutrient name, and nutrient amount

144
00:07:24,001 --> 00:07:26,003
coming from these four tables.

145
00:07:26,003 --> 00:07:29,004
So you can see that this is the product name,

146
00:07:29,004 --> 00:07:31,001
this is the brand owner,

147
00:07:31,001 --> 00:07:33,004
and then all these nutrients.

148
00:07:33,004 --> 00:07:36,007
One thing you notice, some nutrients have zero amount,

149
00:07:36,007 --> 00:07:39,003
which we will ignore when we get to coding.

150
00:07:39,003 --> 00:07:42,001
So these are the four tables that we will work with

151
00:07:42,001 --> 00:07:45,000
in our initial cycle of development.

