Monday, June 16, 2014

VMware Office of the CTO, High Performance Computing

If you know me, then you know that I have a passion for HPC and virtualization and I enjoy a challenge.  I have been working with several customers behind the scenes on virtualizing HPC and developing this market for VMware and had presented at VMworld last year with UCSF on early work virtualizing their genome pipelines.  Now my latest challenge is to join full-time with VMware's Office of the CTO and I am humbled to be able to learn and contribute here.

I will be working with Josh Simons, leading HPC for VMware's OCTO and an HPC veteran and formerly with Sun Microsystems, towards advancement of how VMware approaches this market and the unique problems inherent in HPC as well as the problems now being shared by large "Web-scale" distributed systems.  It was at Hadoop Summit 2 weeks ago that I saw a lot of parallels resolving between HPC and the Hadoop ecosystem.  The common problems are classic computer science issues such as resource management and utilization as well as scheduling at different layers of the system.

And of course all of this is moving so fast, it's hard not to get distracted.  Which is why I appreciate being able to focus on this space because I will be able to leverage my background in compute, network, and storage as well as work on the bleeding edge on integrating and optimizing those with next-generation applications and frameworks.

Thursday, May 29, 2014

Quick Random Question #2

In a high latency, low bandwidth ROBO deployment scenario, how can a customer optimize their ESXi image for deployment so that the remote site will not suffer due to VIB deployment, function independently and allow for unreliable connectivity?

For starters:
1. On a vanilla ESX image, run "esxcli software vib list" to get the baseline of included vibs.
2.  On one of their prod ESX images (clustered, updated, drivers, everything to match what a prod host will look like at the remote site), run "esxcli software vib list" to get all the vibs they'll need.
3. Use ESXi image builder to create the software depot you need and they can export an iso to boostrap any remote hosts.






Additional links
Installing patches on an ESXi 5.x host from the command line
http://blogs.vmware.com/vsphere/2012/04/using-the-vsphere-esxi-image-builder-cli.html

Hadoop Summit and Hadoop as a Service

In the beginning of this year, I mentioned working on Big Data and virtualization and it has been a fruitful time.  Next week I will be co-presenting with Chris Mutchler from Adobe on "Hadoop-as-a-Service for Lifecycle Management Simplicity" at the Hadoop Summit conference in San Jose, CA.  Our session will be on Wednesday from 4:35pm-5:15pm.

I am humbled and excited to help present alongside other sessions from some of the most respected names in the industry from Yahoo!, Google, Cloudera, Hortonworks, MapR, Microsoft.  The growing depth, evolution, and community of the Big Data ecosystem is impressive, to say the least.  I hope to attend other Hadoop customer sessions as well as investigate what other large players are accomplishing from their respective stacks.  I see a lot of advanced sessions around new use-cases for Hadoop and research of adding additional layers and abstractions to Hadoop.  The Adobe session is focused on the usability of Hadoop from an IT operations perspective with a few key points to make:
  • Explain why virtualizing Hadoop is good from a business, techincal and operational perspective
  • Accommodate the evolution and diversity of Big Data solutions
  • Simplify the lifecycle deployment of these layers for engineering and operations.
  • Create Hadoop-as-a-Service with VMware vSphere, Big Data Extensions, and vCloud Automation Center
Hadoop is truly becoming a complicated stack at this point.  After I started actually getting hands on and working with customers on Hadoop-specific projects in 2011, I found that calling this new technology Hadoop seemed a bit disingenuous.  There was really MapReduce and HDFS, a compute layer and a storage layer.  Even though they were tightly coupled, that was enforced for very good and simple reasons.  Spending even more time on this has given more perspective on the different layers and their corresponding workloads.  Unless you're only running one type of job for your compute layer and sucking in data from a static set of sources for your data layer, then these workloads will vary as well as vary independently.  However, in a physical world, with both layers exactly coupled, how can they scale independently and flexibly?

Enter virtualization and everything I've been working on around virtualizing distributed systems, data analytics, Hadoop, and so forth.  Consider the layering of functionality for different distributions and look for the similarities.  If you take a look at Cloudera:


Or Hortonworks:
And Pivotal HD:

As a wise donkey once talked about in a movie when describing onions and cakes, they all have layers and so does any next-gen analytics platform.  Now we have our data layer, and then a scheduling layer, then on top of that we can look at batch jobs, SQL jobs, streaming, machine learning, etc.  Many moving parts and each one with probable workload variability per application, per customer.  What abstraction layer helps pool resources, dynamically move components for elastic scale-in and scale-out and allows for flexible deployment of these many moving parts?  Virtualization is a good answer, but also one of the first questions I get is "How's performance?"  Well, I have seen vSphere scale and perform next to baremetal.  Listed below is the link to the performance whitepaper detailing performance recommendations that have been tested on large clusters.

Speaking of all these layers, this leads to complexity very quickly so another angle specifically to the Adobe Hadoop Summit presentation is around hiding this complexity from the end-developers and making it easier and faster to develop, prototype, and release their analytics features into production.  Some sessions are exploring even deeper, more complex uses of Hadoop and I am eager to see their results, however, enabling this lifecycle management for ops is essential to adoption of the latest functionality of any vendors' Big Data stack.  VMware's Big Data Extensions, and in this case with vCloud Automation Center, allows for self-service consumption and easier experimentation.  There's a (disputed) quote that has been attributed to Einstein that states "Everything should be made as simple as possible, but not simpler."  There are a few vendors working on making Hadoop easier to consume and I would argue simplifying consumption of this technology is a worthwhile goal for the community.  Dare I say even Microsoft's vision of allowing Big Data analysis via Excel is actually very intriguing if they can make it that simple to consume.

Another common question I get is "Virtualization is fine for dev/test, but how is it for production?"  First, simplicity, elasticity, and flexibility are even more important to a production environment.  And maybe more importantly, let's not discount the importance of experimentation to any software development lifecycle.  As much as Hadoop enterprise vendors would like to make any Hadoop integration turnkey with any data source, any platform, any applications, I would argue we have a long way to go.  Any innovation depends on experimentation and the ability to test out new algorithms, replacing layers of the stack, evaluating and isolating different variables in this distributed system.

One more assumption that keeps coming up is the perception that 100% utilization on a Hadoop cluster equals a high degree of efficiency.  I am not a Java guru or an expert Hadoop programmer by any means, but if you think about it, it would be very easy for me to write something that drives a Yahoo! scale set of MapReduce nodes to 100% utilization but which really gives me no benefit whatsoever.  Now take that a step further as that job can have some benefit to the user, but still be very resource inefficient.  Quantifying that is worthy of more research but for now, optimizing the efficiency of any type of job or application specification will allow better business and operational intelligence to an organization and actually make their data lake (pond, ocean, deep murky loch?) worth the money.

Add to these business and operational justifications the added security posture:
http://virtual-hiking.blogspot.com/2014/04/new-roles-and-security-in-virtualized.html
and now you should have a much better idea of the solutions that forward-thinking customers are adopting to weaponize their in-house and myriad vendor analytics platforms.

Really exciting tech and hope to see you next week in San Jose!


Additional links
Hadoop Summit:
http://hadoopsummit.org/san-jose/schedule/
http://hadoopsummit.org/san-jose/speakers/#andrew-nelson
http://hadoopsummit.org/san-jose/speakers/#chris-mutchler
Hadoop performance case study and recommendations for vSphere:
http://blogs.vmware.com/vsphere/2013/05/proving-performance-hadoop-on-vsphere-a-bare-metal-comparison.html
http://www.vmware.com/files/pdf/techpaper/hadoop-vsphere51-32hosts.pdf
Open source Project Serengeti for Hadoop automated deployment on vSphere:
http://www.projectserengeti.org/
vSphere Big Data Extensions product page:
http://www.vmware.com/products/vsphere/features-big-data
How to set up Big Data Extensions workflows through vCloud Automation Center v6.0:
https://solutionexchange.vmware.com/store/products/hadoop-as-a-service-vmware-vcloud-automation-center-and-big-data-extension#.U4broC8Wetg
Big Data Extensions setup on vSphere:
https://www.youtube.com/watch?v=KMG1QlS6yag
HBASE cluster setup with Big Data Extensions:
https://www.youtube.com/watch?v=LwcM5GQSFVY
Big Data Extensions with Isilon:
https://www.youtube.com/watch?v=FL_PXZJZUYg
Elastic Hadoop on vSphere:
https://www.youtube.com/watch?v=dh0rvwXZmJ0

Tuesday, May 6, 2014

Virtualized HPC and Customer Sessions For Your Consideration...VMworld 2014

Now that the floodgates are open, I wanted to let you know about some key sessions regarding virtualized high performance computing applications, analytics and Big Data.
https://vmworld2014.activeevents.com/scheduler/publicVoting.do

Session 1682: Agile HPC-as-a-Service with VMware vCloud Automation Center
What I am aiming at is bringing the simplicity of deployment and integration that is available with Big Data Extensions to HPC clusters.  A simple idea that's already gotten very complicated very quickly but worth the effort to research.

Session 1688: Is Someone Bitcoin Mining on my Cluster? High Performance Security for Virtualized High Performance Computing
If we have these targeted compute clusters, what are we doing to make sure they are being used appropriately?  This issue will only become more prevalent so why not learn how to get in front of it?

Session 1428: Hadoop as a Service: Utilizing VMware vCloud Automation Center and Big Data Extensions at Adobe
Real world details of a customer I have worked with virtualizing and automating Hadoop deployments, from Chris Mutchler of Adobe.  This will be focused on the automation and self-service flexibility gained through virtualization and leveraging BDE.

Session 1424: Massively scaled VSAN Implementations
Another real customer implementation detailed with Frans Van Rooyen, a Compute Platform Architect at Adobe, around their use-case for VSAN for large-scaled analytics.

Session 1466: High-Performance Computing in the Virtualized Datacenter
Edmond DeMattia from the JHU Applied Physics Laboratory discusses his latest real world experience at scale of pooling virtualized compute clusters.  His session was a hit last year and I hope he gets a chance to update everyone this year on his work.

Session 1856: How to Engage with Your Engineering, Science, and Research Groups About Virtualization and Cloud Computing
This session is being done by my two good friends Matt Herreras, systems engineering manager for SLED and Josh Simons who works for the Office of the CTO focused on HPC.  For virtualizing HPC to work, getting the buy-in from the end-user is definitely necessary.

Session 2508: Extreme Computing on vSphere
Both Josh Simons and Bhavesh Davda from the Office of the CTO at VMware presenting on their virtualization of latency-sensitive workloads using Infiniband, GPGPUs and Xeon Phi from virtual machines.

Session 1539: Why the Hypervisor isn't a Commodity. Performance Best Practices from 3 Tier to HPC workloads
Bhavesh Davda and Aaron Blasius, a Product Line Manager for vSphere, discuss performance tips and tricks from the VMKernel.

Session 1232: Reference Architectures and Best Practices for Hadoop on vSphere
Justin Murray from VMware Tech Marketing and Chris Greer from Fedex discuss their current architecture for Hadoop on vSphere as well as look at how this is evolving for the next generation of Hadoop.

Session 1697: Hadoop on vSphere for the Enterprise
Joe Russell's, PM of Storage and Big Data for VMware, first customer panel discussing their experiences of virtualizing Hadoop including Northrop Grumman, Fedex, and Adobe.

Session 1807: Best Practices of Virtualizing Hadoop on vSphere - Customer Panel
Joe Russell's second customer panel including Adobe, Wells Fargo, and GE.

Customer Sessions:
Also, in addition to the customer panels, Adobe and JHU APL sessions, I would highly recommend voting for customer sessions on the given technologies that you want to see.  I always get asked for customer references for all different kinds of technologies and now is the opportunity for you to invite customers who want to talk about their implementations and lessons learned.  Please take advantage and help validate the work that these customers have done to advance the community.  A few examples in no particular order and certainly not a definitive list:
1400, Kroger and their ROBO implementation
2382, 2505, Symantec and their cloud implementation
2463, Francis Drilling Fluids and their SDDC
2770, University of Wisconsin and migrating to the vCenter Server Appliance
1526, Greenpages/LogisticsOne and VSAN
1635, MolsonCoors and their SDDC, including virtualizing SAP, using vCOps, VIN, and SRM
1897, Boeing and their ITaaS
2687, McAfee and Intel elastic cloud
2385, McKesson OneCloud
2285, Grizzly Oil and Horizon View

Thanks for taking the time to read and vote!

Friday, April 25, 2014

Much Ado About Containers


There has been a lot of sudden interest in containers, focused in tech publications fervently discussing “Docker vs virtualization”, or around the Red Hat Summit, blogs and so on.  In my opinion, they are leaving out two words, "Cloud Foundry".  Through the buzzword bingo, it would appear that the Redhat/OpenShift camp, a Cloud Foundry competitor, is aligning with Docker.  To contrast, Cloud Foundry has used Warden for the "container" form factor for building and linking cloud applications.  Since being put into the community's hands, versus VMware or Pivotal's, it's possible that Docker will even become a choice here as well (Decker).  Ultimately, customers’ options haven't really changed even though many articles portray that a turning point is imminent.  Customers will be able to utilize the virtual form factor that fits their business needs, and dare I say fits their business culture?  

This could be on a VMware SDDC, or PaaS, which can certainly be deployed on VMware-based IaaS or of course other alternatives.  It depends on what set of abstractions they are comfortable with even though any startup focused on Docker needs to swing the needle to their side in order to better justify their existence.  I am a fan of Cloud Foundry and PaaS in general, but I don’t believe that every app will be run at that layer of abstraction.  But hey, maybe that last statement will be one of those quotes like “640K ought to be enough for anybody.”  I am just thankful to be working with such a diverse group of customers to keep this in perspective as well as get to play at the bleeding/leading edge.

Additional links:
https://groups.google.com/a/cloudfoundry.org/forum/#!topic/vcap-dev/V-lVpMpNqL4/discussion
https://docs.google.com/document/d/1DDBJlLJ7rrsM1J54MBldgQhrJdPS_xpc9zPdtuqHCTI/edit
http://blog.cloudfoundry.org/2013/09/08/combining-voice-with-velocity-thru-the-cloud-foundry-community-advisory-board/

Tuesday, April 8, 2014

New Role and Security in a Virtualized Environment

I've accepted additional responsibility in VMware’s CTO Ambassador program so in addition to working with the Office of the CTO on advanced projects around vHPC and vHadoop I will be working to get more visibility and feedback from customers directly to the PMs and R&D engineers responsible for our products.

Despite not officially being in VMware’s Network and Security business unit (NSBU) for the past year, I am still constantly called in to customers to discuss security in a virtualized environment.  In the past, I was responsible for establishing security documentation and baselines for many military branches such as the USMC and how they could achieve their security accreditation for any given site.  This effort to be able to apply DISA STIGs to virtual environments predated VMware’s standard security hardening guidelines.

Now, all of VMware’s security information can be found at http://www.vmware.com/security.  Customers want security reference architectures but they also want more detail about how to customize those for their specific environment.  If you take a vanilla installation of vSphere and apply the latest security hardening guide, you have the basis for the reference configuration that was submitted for Common Criteria certification.  For example, vCloud Networking and Security (vCNS) v5.5 recently achieved EAL 4+ certification.

Two of the biggest issues I still see in virtualized environments are inconsistency and complexity, which have a direct bearing on companies’ security posture.  One would think that with standardized templates, tools and scripts that consistency would be much easier.  However, production IT departments have to support a wide range of applications and this can quickly devolve into catalog sprawl and managing a fairly complex array of templates.  In addition, security policy becomes increasingly specific to each VM or application and again too complex to manage.

I believe the idealized notion of hackers sitting in front of their keyboard elegantly dissecting a target environment is more than a bit disingenuous.  Malware toolkits, or “Sploits”, built to do all of the tedious work are getting even better at spotting inconsistencies and exploiting them before the IT and security admins who should be most familiar with a given environment.  We used to make fun of “script kiddies”, those who ran scripted attacks without understanding why they would be effective, but the threats have evolved as always.  In an autonomous fashion, attacker toolkits can explore a local OS and network, probe for weaknesses, and evade detection in running memory.  One of their typical tasks besides evasion is to compromise any local security controls, antivirus, etc.  So why not only base security policy in the network at a different layer?  Those local controls do provide the best amount of context to what is actually going on within the OS, but without appropriate isolation, they can be easily bypassed.  Why should the IT and security admin toolkit be any less dynamic and automated?

The “Goldilocks Zone” first discussed by NASA and reinterpreted for security by Martin Casado of VMware and Nicira fame keys into leveraging virtualization as being “just right” to support security in a virtualized environment.  To me, and many customers I’ve spoken with, the virtualization layer, along with providing the baremetal abstractions, delivers the right amount of context for local apps and OS while isolated from potential malware that would otherwise corrupt or compromise those local controls.  ESXi can be this trusted layer, with TXT, and appropriate application of the security hardening guide recommendations.  This can also be audited via 3rd party tools such as Hytrust.

Within vCNS, you can configure security groups to be the logical container for policies.  You can start with static security groups that are defined pertaining to the standard virtual infrastructure layout that vSphere admins are already familiar with such as at the virtual datacenter, clusters, or port group.  The next step is to allow security groups to be dynamic by leveraging the object model AND enable partner solutions to interact with those objects to affect security policy in real-time and coordinate security response to violations.  NSX allows this through essentially creating security tags for VMs that may be read by NSX as well as NSX ready partners, for example Trend Micro and Palo Alto Networks.

Much more to be said about this, but needless to say, excited to be working more directly with customers on next-gen apps with NSX.  I will actually be discussing this live with Trend Micro and Accuvant on 4/18/14 and will post a link to a recording when it is available.

Additional links:
http://blogs.vmware.com/networkvirtualization/2014/03/goldilocks-zone-security-sddc.html
http://www.vmware.com/security/certifications
http://www.forrester.com/No+More+Chewy+Centers+Introducing+The+Zero+Trust+Model+Of+Information+Security/fulltext/-/E-RES56682
http://www.vmware.com/products/nsx

Tuesday, March 25, 2014

Do not forget NTP

I have seen a lot of recent issues related to NTP and time skew in VMware environments recently.  Any new appliances which have SSO as a dependency need time skew to be kept at a minimum.  If you're deploying vCloud Automation Center (vCAC) 6.x, then time needs to be inline with some consistent source, most likely the domain controller.  This also applies for the VMware Big Data Extensions (BDE) appliance and if you've deployed the vCloud Networking and Security (vCNS) Manager aka vShield Manager appliance in the past with vSphere 5.1. The errors may not be particular helpful unless you dig in the logs and see something similar to the following "Server returned 'request expired' less than 0 seconds after request was issued".

In a Windows world, the servers can be pointed at the Domain Controller and on a linux distro you typically point /etc/ntp.conf to one of the pool.ntp.org servers and always make sure UDP port 123 is open.  In a virtualized context, you could be lazy and synchronize with host time from VMware Tools.  Most of the appliances now have this exposed under the "Admin" tab.  But let's say you go through the ESXi host configuration and point the NTP client to the domain controller.  That should fix it right?

If you're still running into the issue and noticing that your ESXi hosts are not syncing, then you need to read this:
Synchronizing ESXi/ESX time with a Microsoft Domain Controller
because there are several steps that require going into the ESXi shell to remedy.

In the interest of strong design, please don't take NTP for granted.  Your logging, which can be mind-numbing to troubleshoot to begin with, obviously makes no sense if there are time discrepancies.  Then we have log management tools like Log Insight now for example, which have basically become mandatory.  Do you still set all your servers to a specific timezone or have you standardized on UTC?


Additional Links:
http://tycho.usno.navy.mil/NTP/
https://blogs.vmware.com/management/2014/01/vmware-vcenter-log-insight-1-5-technology-and-features.html
http://blogs.technet.com/b/askds/archive/2007/10/23/high-accuracy-w32time-requirements.aspx